It really is dumb to be arguing over tabs vs spaces, after all.
It really is dumb to be arguing over tabs vs spaces, after all.
Blank lines is an obvious one: where do you insert them to "group" sections of a 30 line function? Or are there no blank lines at all?
Line length: just "wrap at columns X" (or never wrap) is not enough, because people can and do wrap at specific locations for specific reasons, because that makes more semantic sense or looks nicer than cramming as much as possible.
Whitespace gives readers subtle information about the structure of a function before they read any of it. Its a powerful tool. Dismissing or - worse - deleting whitespace wholesale sounds profoundly misguided to me.
Why make code harder to read? Where's the benefit to your approach? I can't see any.
https://doc.rust-lang.org/src/core/slice/mod.rs.html#2786-28...
The function is pretty short - 40 lines including comments. Despite how short the function is, it still uses whitespace to separate and group adjacent lines of code. Personally, I find the code more readable like this. Indentation makes syntactic blocks obvious (the while loop and if statement). But there are also conceptual groupings between lines that mean nothing to the compiler, but are semantically meaningful to humans. I can tell at a glance that the comment about safety is most associated with that one line below the comment. And so on.
I think this code would be worse if we deleted the whitespace. How would you improve this code?
At a minimum, the blank lines implicitly scope the comments that sit above connected code blocks. Those comments - especially the comment about safety - would be harder to understand and audit without its context being so clear.
And again, can you name any benefit to removing all the blank lines? If we can both read the code easily with some empty lines to space it out, that seems like the best option.
On top of that, now you need to write a parser and compiler for your AST file. It's probably very simple and does rudimentary validation, but that defeats the point of the AST - it's a valid, canonical representation of a program by construction.
All in all, it seems like a good idea, and people have done it. But there's also good reasons to be apprehensive.
And at the end of the day, you need to ingest text as input, and you need to do it as fast as possible. There's not a ton of benefit to keeping an AST around on disk and in sync with the text that generated it when you are already able to compute it faster than you can read and deserialize it.
In an in-house dialect of Haskell I used to work with, we solved this problem by just making tabs a syntax error. Never had any problems.
(I think tabs might be have been allowed inside strings.)
(I don't really know what happened to it since.)
“We’ll just force everyone to do things one way” kind of ignores the point I was making. It shouldn’t be necessary for you to care how anyone else formats their code, same as you don’t care what font their code displays in or what text editor they use. It feels like a vestigial aspect of programming that we have to concern ourselves with it in 2024.
FWIW, just happy to have a chance to unload this thought finally: it had surprisingly little impact on code reviews, in that the "personal preference I need to enforce" just ascended abstraction levels.
And yes, that doesn't help you if ex. your style is a blank line following every code line.
In practice, it works, I surmise because people are fine with someone else's code being in a different style, but they want to write in their style.
Wasn't a problem in practice. Just like we never had any problems with anyone wanting to use eg Pascal or so.
In a language like C your outermost layers of leading whitespace are always indentation, and then you might have some alignment inside.
But in Haskell you might want to align arguments to a function, but some of the arguments can have blocks inside of them.
Mixing up tabs and spaces is technically possible, but it's too much of a pain in practice to bother.
More technically, language servers usually have a CST that they use to build the AST incrementally, and the AST contains references back to the CST that generated it. This is what allows you to handle incremental text edits and compile small deltas to the AST instead of the typical batch compiler design that attempts to parse everything all at once.
I've seen language server that completely ignores the parts with error, and I much prefer error nodes because then I still know there is something and these error node can still have children
There is also the problem that an error returned by a language server is a class of a "diagnostic" that includes syntax errors, semantic errors, warnings, lints, etc, associated with a span in the source code. It's much easier to think of diagnostics as a separate data structure that gets filled up during lexical/semantic analysis and associated with spans in the full syntax tree (you can even store them there as fields, but they don't necessarily have children). Then it's obvious how the structure gets created and fed back to the user.
And finally, the whole point of an AST is to be a valid canonical representation of a program so the compiler query it drives doesn't have to do additional input validation. So it just makes the queries/compiler passes easier to write.
You're right, ex. https://dev.to/jillesvangurp/comment/6gnb/
Sadly there's enough bitrot that ex. intentsoft.com is offline.
More here, but article is v opinionated/judgement oriented, comments are useful, but again, bitrot :( https://wiki.c2.com/?IntentionalProgramming
It strikes me that what I am describing as "bitrot" may also be "never really shipped, so the vagueness isn't accidental"
This always sounds more difficult on paper than just wrestling dependencies till dawn, upgrading from JDK 11 to JDK 17, for example. So I usually give up the mental exercise there.
Plus, following a file of transforms is mind-bending: someone may follow a method definition with a pattern seek, and then start appending some more code. Context is lost. It would be literate programming only with enough empathy for comments.
Which is all to say, would it be easier to move between Spring versions if the app's commit history were a series of transforms instead of changes to static files?
Suppose a commit establishes a framework version, and then follow a bunch of commits for domain objects, a skeleton controller, and so on. If we could play those decisions forward, but edit the transform instead of the source, would it be easier to dissect which next dependency to manage?
This loops back to ASTs: we would still edit and change files, but the history would be ed(1) macros (or something better, like ASTs). Somehow, it feels like there could be reconciliation between source control and "manipulating a timeline of changes."
Git may already have this, or a simple while loop with some decisions about how far to play the changes, like editing a cassette tape. A list of patches to apply, with pre- and post- hooks for rules scripts.
This saves you the trouble of authoring a separate programming language, or finding a way to preserve all of the niceties of the original syntax and formatting that wouldn't directly translate to an AST (like how many newlines are after a particular stanza or function).
Case in point: Recast (https://github.com/benjamn/recast) is of particular interest with regards to JS/TS in this vein, because it does preserve a lot of the spirit of the source in its conception of the AST. But also last time I used it (couple years ago now) it would explode on any code with an emoji in it. It's genuinely not an easy problem.