Better Clojure formatting
tonsky.me
tonsky.me
Furthermore, I personally find a lot easier to read code that is not too spread horizontally. I do not have extensive experience with Lisp or Clojure, but all languages I've worked with will handle a "small" line width of 80 character just fine most of the time.
I do recognize this is a matter of personal preference, but please don't dismiss the 80-character limit because it's old. It may be, but it is still relevant.
Other than that, I wholeheartedly agree with the post. I started toying around with clojurescript for the past few months, and the one thing I can't wrap my head around is its formatting. Sometimes one space, sometimes two. Sometimes block spacing. It's very confusing.
In Clojure code, the most likely reason to go over 80 characters for me is let forms like (let [long-descriptive-names-that-get-too-close-to-80 (and we-have-hit the-limit)]) but overly long names mess the code up anyway and should be avoided. Descriptive is good, but too long makes it hard too read too.
Changing your standard to a value so close to the existing standard seems like an odd choice are you using a verbose language?
Guessing you don't write Java? While you could do it, you'd have so much line wrapping that I imagine it would be much more annoying/difficult to read.
The counterargument is that they suck at it, you get a break (maybe word-aligned) with an arbitrarily indented next line.
As a purist I advocate for a future with smarter editor soft wrapping.
What would be useful would be a text editing mode that automatically scrolls horizontally to keep the code in view, regardless of indentation amount. Then you can have your 80 column window.
What sucks is moving down in a body of code, and watching the text disappear out to the right because the cursor is tracking straight down through the indentation instead of riding to the right of it.
And I'm guessing it would only worsen the problem of sprawling code that's difficult to read, by removing one of the few sources of pain that encourages developers who are less mindful of what it's going to be like for the person who'll have to read this stuff 3 years down the road to exercise at least a little bit of care.
(.....
(....
(...
(.........
(..
(...)
(....)
(...(...
(.....)))
( )))
(............))))...)
If the 80x50 peephole moves nicely along the diagonal, I'm okay.I don't want to go from this:
+-----------+
|(..... | ;; not to scale, obviously
| (.... |
| (...) |
| (.....|...
+-----------+.
(...)
(....)
(...(...
(.....)))
( )))
(............))))...)
To this: (.....
(....
(...
(.........
(..
+-----------+...)
| |....)
| | (...(...
| | (.....)))
| | ( )))
+-----------+............))))...)
When I move the cursor down. If the box would move diagonally, that would be great. Yes, of course the box will have empty space at the bottom left and top right.That could be addressed in a funky way: why not rotate the text about 45 degrees? Then the rectangular window would capture a fuller slice of the diagonal. I could get used to the rotated text.
Whether it's single lines that make a beeline for the pure land in the east, or towering ziggurats of stacked contexts slowly plodding their way into the sky, it doesn't really matter all that much to me. I count both of them as "sprawling code that's difficult to read." Lisp's a functional language; it's totally acceptable to break things up with a `defun` (or I guess, in this case, `defn`) here and there.
> This is one of the stupidest things to automate. Unlike indenting, removing trailing whitespace simply produces diffs where it doesn’t matter. And does nothing more.
I suppose formatting is a weak point for version control systems, but I think that is all the more reason to get rid of trailing whitespace. The more your code is clean of ambiguous formatting, the less it will change between commits.
If you need to clean up a whole codebase you can throw in something like find -name * .java | xargs sed -i 's/\s* $//' and just commit that once.
Sidebar: How do I print an asterisk and a char without adding in whitespace.
It's only ever a mild nuisance (in that scenario or for diffs), but the fact that it's trivial to automate seems to make it a no-brainer to me.
Maybe 80 isn't the right one, but letting code stretch out to ad-infinitum (either with awkward line breaks or having to scroll) is a big ergonomic loss for me personally.
This seems to work well enough for two columns. Useful for editing, but indespnseble for diffs. For reference my 15" Mac has 234 columns at full screen.
It's easy to skip and forget where you were when you look to compare the code to something and then look back.
I also used to not care about side-by-side, but now, I find it too useful. Besides, nobody wants to scroll to read git diffs.
The core problem (sort of created here) is about how lispy Clojure should be. Unlike most programming languages, Lisp code isn't meant to be fully statically analyzable. This affect indentation as well, because the "full picture" of syntax is only available at runtime.
Consider a macro invocation:
(foo bar
(some args)
(some other code))
vs.: (foo bar (some args)
(some other code))
Which indentation style is the correct one? That depends on what "foo" is. If it's a function, then the first one. If it's just a regular macro, it's still the first one. But for some particular macros - like defun in elisp/CL, the valid syle is the second one. You can't, in general[0], know that until runtime. That's why the Lisps usually give you tools for that. In Common Lisp, you have &body, which works like &rest but also hints that the expected indentation is like the 2nd case I shown here. In Clojure, you have metadata - you can attach e.g. {:style/indent 2} to your macro, which signals that desired indentation changes to 2nd form on third argument.My personal opinion - most Lisps are designed for interactive, in-image work. So is Clojure. If you want to have a completely static formatter, go back to coding Ruby or Python.
--
[0] - You can guess it in simple cases, but not when you're dealing with macro-writing macros, i.e. code that writes code that writes code.
There is no correct one. There are aesthetic preferences to make Lisp code readable. Lisp has code formatters for decades and Common Lisp comes with one in the standard (the pretty printer). Tools to format (not just indent) textual Lisp code has also been available for decades, but are less used.
> If you want to have a completely static formatter
The proposal here is to ignore the programming language Clojure and format all code only on the expression level.
(macroname
first-param
third-param
fourth-param
ad-infinitum)
This is the only case (there are no special, or complex, or any other cases).Besides, if you took your time and actually read the article, you'd see `defn` in the examples.
At the syntax level, yes. But indentation is not a syntax-level concern, but a semantic one. Speaking Clojure,
(defn function [arg]
{:some-op arg})
has the exact same syntactic structure to (assoc-in some-map
[path]
{:some-key path})
but it's expected to be indented differently because of the difference in meaning between defn and assoc-in.Lisp allows you to use macros to provide new symbols that, put first in a list, imply something else than function call. So user-defined forms are put on equal footing with things like if or defn. And user-defined things can generate more user-defined things that are meant to be indented like defn. Hence the need for loading the code for correct indentation.
(lispword a b c
d e f
g h)
whereas symbols not in the set like this: (not-lispword a b c
d e f)
This is important so we don't end up with this: (let ((x y))
form
form)
or, vice versa, this: (list (item 1) (item 2)
(item 3) (item 4))No, that's a misconception. Lisp has special forms from day one (1960). On day two (1962), macros were added -> they implement even more syntax.
You could get away with parsing the code statically in the most trivial cases, but to do it even remotely correctly, you'd have to parse all the code, and only then format individual files. A lot of work for no good reason, and it wouldn't save you from macro-writing macros.
In a large Lisp project, this would screw a lot of things up. I don't have relevant Clojure example here, but in Common Lisp, I worked on a pretty big codebase (CLIM2), where trying to indent things without loading the code into a Lisp image would screw up indentation in 50+% of the files - because a lot of frequently-used macros were written by other macros.
In Clojure we format everything with spaces (https://ukupat.github.io/tabs-or-spaces/ - 99% of Clojure code uses 2 spaces). This formatting style is pretty much trivial and all code editors support it.
I'm perfectly fine with Clojure code formatted slightly differently depending on a person who writes it. If you want a BDSM language, code in Golang.
We've conflated how the code is stored with how it looks. We need to separate these. That's one of the problems with equating code with text: they're different things. Code is data (or objects), text is a way to represent code.
It's not just variable casing; it's any situation where one coding style distinguishes two things which look the same in another coding style. And I think there's a lot of such cases.
Really aggressive reformatting introduces git blame horizons that can make code archaeology more annoying.
I wouldn't consider it a reason not to switch to autoformatting, but it might be a reason not to get too wild about changing the conventions.
At least when the person sees my commit in git blame they know to annotate the previous version.
Luckily most people have settled on pep-8, but even there, there are points of contention, such as line length limits, or how spaces won the tabs vs spaces debate (despite that imho tabs are superior as they encode indention cleanly: 1 tab = 1 indent, yet each person can set their editors tab width however they prefer) See, nobody will ever fully agree on formatting.
But that wasn't my point really, which was just that despite its forced indentation, Python still has, in my personal experience, its own formatting debates.
And this is the entire point of the article. He discusses this at length: to reach the universality and speed of a `gofmt`-like formatting tool, we would have to give up on the idea that s-expressions should be indented differently based on their content.
Personally I find this a very tempting tradeoff. You may not, but indenting different s-exps in different ways based on their first element is not "correct," it's just... what most people choose to do.
They are not just s-expressions. They are programs. Just like a C code is not a random string, but a program.
https://clojureverse.org/t/clj-commons-building-a-formatter-...
That, plus the fact i haven't coded on a 4:3 monitor since ~2002 makes me thing this is not helpful. The overuse of newlines in coding standards is... a useless exercise.
Constraining line length is great for readability. The benefits are similar to those gained by limiting prose line length, and sentence length.
Think of a long line of code as having similar traversal characteristics to a linked list. This requires more memory for comprehension, and is hard to search.
Now compare this to a newlined block of code, which resembles a richer data structure that spatially colocates relevant elements.
It's a difference that matters a lot.
Not really; prose doesn't indent with increasing nesting, so line length in prose actually means that many non-whitespace characters to scan.
> Now compare this to a newlined block of code, which resembles a richer data structure that spatially colocates relevant elements.
Nobody writes huge lines in Lisps; but the indentation can get out of hand in large functions. So you have this neat block that spatially colocates relevant elements, but has 150 columns of whitespace on the left, which creates a navigation challenge that is different from the problem of scanning a long line.
If you combat this by simpy constraining the column length, then what happens is this problem
(cramped
expressions
(butting up
to
80
col
limit))
You have little recourse but to break it up into smaller functions (perhaps some of them nested in the parent).Clojure lines increase in length as they increase in nesting, but without increasing nesting, they don't typically get particularly long[1], unless you have very long argument lists or very long names. So, the key is to reduce nesting (which helps comprehension anyway) and Clojure has plenty of tools to help with that: you can use ->, ->>, as->, some->, or you can bind sub-expressions to names using let/when-let/etc. There's no need to split code into smaller functions unless it makes sense to do so (eg for code reuse or semantically).
[1] 2-space indents help keep indents from being crazy wide without crazy nesting. The only place where I often find lines getting too long is let-blocks with aligned expressions and long variable names:
(let [this-name-is-too-long (some-code-here
(uh-oh this is getting long)
... ]
))
Had to close those parens to stop the mental itch!> You have little recourse but to break it up into smaller functions (perhaps some of them nested in the parent).
Generally this type of factoring is a good thing for the codebase, and a very natural part of the development process in Clojure.
My clojure code is more vertically compact (despite sticking to 80-column lines and avoiding deep nesting) than a lot of C++ code I see out there:
if (condition)
{ // wasted line
some = code(here);
} // another wasted line?
// sometimes a blank line here too
if (another condition)
{
...
The Java convention at least wastes a tiny bit less vertical space: if (condition) {
some = code(here);
}
if (another condition) {
...
In Clojure I have the same, but without the trailing closing brace. I insert blank lines if I need to visually break the code up or if its really unrelated from the previous line, but its up to me when to apply it based on when it makes sense to do so, not just because the naming standard says so. (if condition
(some code here))
(if another-condition
...)
In reality, I'd of course use when and if the body is trivial then I may even simply write (when condition trivial-body) as one line. As long as it won't go over 80 characters.