Linus Torvalds adds arbitrary tabs to kernel code
arstechnica.com
arstechnica.com
It's good that Linus is really exercising those 3rd party tools! They should send some money his way for helping them test their code.
This web view renders tabs as spaces, so it's not possible so see what was changed.
https://github.com/torvalds/linux/commit/d5cf50dafc9dd5faa1e...
https://github.com/torvalds/linux/blob/d5cf50dafc9dd5faa1e61...
Unfortunately Github doesn't have a way to render symbols for whitespace, but you can tell by selecting the spaces that the previous version had leading tabs. Linus changed it so that the tokens `default` and the number e.g. `12` are also separated by a tab. This is tricky, because the token "default" is seven characters, it will always give this added tab a width of 1 char which makes it always layout the same as if it were a space no matter if you use tab widths of 1, 2, 4, or 8.
https://github.com/Reference-LAPACK/lapack/blob/master/SRC/c...
There's also Elastic Tabstops if you want to go all in on using tabs for alignment: https://nick-gravgaard.com/elastic-tabstops/
The limiting factor on line length should be "is this too complex for your coworkers to understand"
TBH my money is on that never happening, but maybe we will skip to an even higher level, where the canonical format for source code is the AST rather than plain text - indentation, braces, line length, all the things that humans care about but the computer doesn’t become render settings.
Sadly, nobody wanted it. Check out https://mbeddr.com and weep for what could have been.
Notably, this includes CLI shells, connected to a "terminal emulator", where what is being emulated is an ancient piece of hardware:
https://en.wikipedia.org/wiki/Teletype_Model_33
Semantics of the ASCII tab byte code:
https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1...
A far-downstream consequence of this is that source code formatted to an assumption of tab stops at other than 8-column intervals, as is not uncommon in Javascript, produces unreadable CLI output from diff, git-diff, ...
Just that the ingested data is a part of the Linux kernel codebase. Quite some hubris to proceed and apply such a "fix" by making a commit to the Linux kernel...
But I totally sympathise with Linus' annoyance when the issue is with an external tool, and the author didn't explain which tool or give any reason why it's hard to fix that tool.
Unfortunately until somebody spends the time to create a Kconfig test suite, the kernel itself needs to be the test case for this oddity.
Developers change things for parsers all the time. That’s, like, coding.
It's both Linus's plain arbitrary right, and his plain job, his defined role and office, to make exactly such decisions for this project. That's not hubris. It's just a role that affects a lot of people.
What makes it hubris on one side is "Who do you think you are making such a change to the Linux kernel that everyone else will have to accept?"
The reason things like that are phrased as questions is to allow for the possibility that there might be an answer.
For one of these parties, the question is rhetorical.
For one of these parties, the question is not rhetorical.
And THEN not even providing more justification for this in the description.
I'm not claiming anywhere that Linus is the good guy. Or bad guy.
In this case I agree with his stance that whatever this parser is, it should better fail harder in order to get it fixed.
And if you know how he reacts when he's making a big deal of something, you know that this one isn't one of those times...
1. Use all spaces
2. Use tabs for indentation and spaces for alignment
Unfortunately, only individual developers seem competent on their own to do #2, so everyone who cares about readability inevitably practices #1 by default.
You can never use only tabs.
I'm not surprised that this isn't something that projects have been able to adopt successfully very often because I've never found it very intuitive that those are separate things. In what way is "indentation" not also a form of "alignment"?
I don’t mind if people use tabs but mixing the two is not great.
I'm practically unable to `git add` unformatted Go or Rust code, and very happy with the setup:
.config/git/attributes:
*.go filter=gofmt
*.cgo filter=gofmt
*.rs filter=rustfmt
.config/git/config: [filter "gofmt"]
clean = "gofmt"
[filter "rustfmt"]
clean = "rustfmt --edition=2021"Knowing whether a space in a space-formatted document is for alignment or indentation is AI complete.
If monospace was actually a good idea and not a technical limitation of computers from the 80s, books, newspapers, and your comment on this website would be displayed in monospace.
myVars = {
first: 123,
second: 234,
afterward: 345,
also: 456,
} //~~~~~~~~ <- alignment
// or:
def my_long_function(arg1, arg2,
more_args,
and_some_more)
//~~~~~~~~~~~~~~~~~~ <- alignment
This is, e.g., what Go uses for struct fields, and what some Python style guides use for hanging function definitions. Regardless of tabs/spaces preference, both of these are independently bad because they churn diffs unnecessarily: if you change `afterward` to `afterward2` then you need to change all the nearby lines, and likewise if you change `my_long_function` to `my_longer_function`. Some formatters, like Black, Prettier, and (mostly) Rustfmt, avoid this pattern entirely, and they are better for it.* You can do this and still use spaces if you prefer, too.
The problem being referred to here is that a change on one line causes changes to unrelated lines. It makes it harder to pinpoint which line was changed intentionally in a diff and which ones were just caused by tab issues.
If you format before every commit, then the only changes you get are the changes you made.
It is perfectly doable to do only tabs, but many end up mixing in spaces.
The curse of space-only files is in people that manage commit indentation errors, breaking auto-detection in some editors, which propagate to even more indentation errors... All it takes is an inattentive reviewer, or review-less merge.
Sure, but could i not say the same of using any indentation at all?
Not to mention that in languages like Python, indentation is syntax.
a_variable = (
'lorum ipsum dolor sit amet ' +
'my poor memory has left me quite upset ' +
'for i cannot remember what word comes next '
'in this long descriptive text' +
'surely this is bound for the incinerator ' +
'but remember any haiku can end, refrigerator.'
)
But now if I choose to align certain characters: a_variable = (
'lorum ipsum dolor sit amet'
+ ' my poor memory has left me quite upset'
* ' for i cannot remember what word comes next'
' in this long descriptive text'
+ 'surely this is bound for the incinerator '
+ ' but remember any haiku can end, refrigerator.'
)
… the errors in the first version are now plainly obvious. (Both the missed space, as well as the missed +.)(This is an example. Yes, there are languages for which you don't need the +. There are some for which you do, however. There are also some that resist having the + moved about: for example, in Javascript, the parens become required, or you'll trigger the horrid auto-semicolon "feature".)
Either you misunderstood what we mean by alignment (single whitespace separating operators and operands are not alignment), or maybe you tried to indent the first string literal further to align the string literal start boundaries. I sincerely hope it wasn't the latter...
Putting the operator on the second line, at the same indentation level, is a perfectly fine, indentation- and alignment-agnostic stylistic choice, and has its uses.
I have changed the alignment of both ' ' and '+' characters between the two examples, from not being aligned to being aligned.
> as all your lines of code start on the same integer indentation boundary
Indentation is not the only thing "style" guides enforce.
> and changed the bug between the two examples.
The bug should be the same in both examples, though I do see I transposed a '+' into an '*'. That wasn't intentional; the two are supposed to form the same AST, but with purposeful alignment in the second making bugs in both far more visible.
(You seem to be using "alignment of characters" to mean "alignment with indentation", which is more limiting than I would take "Aligning a character on one line with an arbitrary character on another line" to be.)
- If you have project wide automatic code formatting: Tabs
- Otherwise: Spaces
Nowadays, most of my projects use option 1.
What do you mean?
> You can never use only tabs.
You can.
The pains is, most website or editor never handled that well enough. You end up have mixed tab/space at unexpected position and never knew about it.
Just banning the tab is probably not the most 'correct' option to fix it. But it is the most feasible one to get the job done. Because fixing all the tool, editors and websites is nearly impossible for an average man.
Reformat-on-every-commit really only works for highly-centralized, tightly-coupled, monorepo-using monolithic organizations. Basically the exact opposite of kerneldev. For those folks reformat-on-every-commit works great.
Realistically, either of these options would only really work for tools and codebases designed with these workflows in mind, as they'd be useless for inputs to inherently text-based tools like the C preprocessor.
That's the easy case.
The hard case is situations where the formatting (whitespace, etc) is not simply some deterministic function applied to the AST. For example: comment placement and tabular layout of repetitive constant definitions. Or situations where vertical alignment is used to call out parallel structure.
Mathematicians have been using horizontal alignment to call out important structure in equations (both written and typeset) for hundreds of years. Here's are two examples if you don't know what I'm talking about:
https://kagi.com/proxy/ILVap.jpg?c=085W0W0Vgv9kix8Uo-TAMYM_F...
https://kagi.com/proxy/Zr5re.png?c=085W0W0Vgv9kix8Uo-TAMb-ZJ...
The people who stubbornly insist that programmers have no need for this sort of expressiveness seem like innumerate cavemen, frankly. They cheapen our field, reducing it from informatics to mere keyboard labor.
The urge to run code formatters is a "language smell" indicating that your language has way, way, way too much odd and irregular syntax. See Haskell for a counterexample. Sure, there are a few code formatters for Haskell, but nobody really uses them. The language's syntax is so minimal and clean that there just aren't all that many preference-based choices to be made.
I haven't run a survey, but I get the impression that code formatters are widely used in the Haskell world, especially Ormolu and Fourmolu.
Ormolu and Fourmolu
That's only one formatter; the latter is a fork of the former... which is the in-house formatting tool of one particular consulting firm. It has only existed since 2018, and is said to be mainly a response to golang/rust people trying to impose their culture on the Haskell world.
Pointing at ormolu is kinda like saying that formatters are popular in the Java world because google-java-format exists. The fact that that project exists says more about the company that created it than the community around the language it operates on.
Ah, as a fact it presumably has some form of evidence that demonstrates it. Would you mind sharing that evidence?
These already exist: https://github.com/Wilfred/difftastic
when we get that, then we should get even less merge conflicts.
Counterintuitively, that is not the case. AST-merge is a much, much, much, much, much harder problem than AST-diff.
https://github.com/Wilfred/difftastic?tab=readme-ov-file#can...
The fact that diffs can be used to drive a 3-way merge is in fact an accidental property that arises due to the sheer crudeness of the line-based diff format. As soon as you start using more-sophisticated diff formats, solutions to "the diff problem" no longer lead directly to solutions to "the merge problem".
But on a more serious note, in my experience I've not had any issues with Go or Rust codebases (for example). Not using their formatters is heavily frown upon, so I haven't really seen any reformat happen at all; not in my bubble at least[1].
Other languages, on the other hand? Yeah, good luck with trying to have consistent formatting. Even if a project has formatting rules "enforced", there's always (always) going to be an exception, bikeshedding, etc.
[1]: Unless it's someone obviously very junior. The few times I've noticed badly formatted code in Go, has been in random repos from someone who clearly didn't have that much programming experience in general (looking at how code was written).
"Our Python codebase is old af and has inconsistent styling. Should we start using black/ruff to format it?"
"Hmm that's maybe a bit too intrusive. Could autopep8 be enough? Since it follows the recommendations of the PSF, so it's a more 'official' way of formatting, and it should be more future-proof [famous last words]."
"Hey, I heard this pyxkcd927 formatter just came out, and it also automatically fixes linter warnings [meaning, different output]. Any thoughts on migrating to this?"
In general, any software engineering thing that causes multiplicative work for the rest of the org should be handled this way.
For instance, at a prior job, it cost 100+ engineers a week or so of productivity to rename master to main because tooling and automation had hardcoded master all over the place.
It cost the person that initiated the change an hour or so, and I’m sure it helped their promotion case.
Of course, accounting didn’t compute the cost of this to the business ahead of time. If they had, I’d hope they would have insisted we spend the money on recruiting and hiring more diverse candidates instead.
Newbies editing code is a problem best solved with automatic formatting.
Type myvar = something_very_long +
that_needs_alignment;
The only way to guarantee such complex alignment is using spaces. It renders everywhere exactly the same. So reviewing stuff on web-based environments gets easier.It's just that sooner or later someone will get it wrong.
https://nick-gravgaard.com/elastic-tabstops/
I'll try to do my part.
I've been writing (and reading) a lot of Lisp lately, and have seen some projects adopt indentation idioms which offset blocks from the regular tab-stop intervals. A standard tab-stop might be +4, but then arguments might be lined up with each-other at +6 (+1.5 stops):
(list 'a
'b
'c)
Any of those arguments could be forms, which might return to the standard tap-stop of +4 but off by 2 (let's call that `+4(+2)`): (list 'a
(with-foo quz
(bar quz))
'c)
Lists written with the `quote` syntax might do the same thing, but produce an offset of only 1: '(a
(b
(c)))
The end result is that changing a code's indentation level (like when raising an anonymous function into a top-level definition) means adjusting the spaces (which were supposed to only be for alignment) if any preceding forms had introduced an offset. The alignment doesn't change, but it's relation to regular tab-stops does. Sometimes, this is an adjustment of zero, but you still need to think about it. In an all-spaces environment you could use vim's visual mode to select the block as a rectangle and paste it into the top level, but with "indent/align" semantics this same situation regularly requires maintenance (by inserting or removing spaces, in addition to manipulating indentation via tabs) to preserve both proper alignment and the integrity of the semantic distinction[1].I think that elastic tab-stops fix this by correctly modeling indentation is it should be conceptualized, platonically divorced from the characters encoding it. I think that spaces are better than tabs when white-space isn't significant (and personal style might disregard tab-stops), but recognize that tabs are more accessible and worth using when tab-stops are enforced as a byproduct of significant white-space. I repeat that elastic tab-stops are awesome.
1: Someone correct me if vim can actually just do that with tabs just fine, i dont really know, it late~
This is no worse than using 4 spaces to indent. Using 2 or 8 or 3 spaces for a tab is a skill issue.
It’s perfectly fine to still have a canonical width if you want to use it for other reasons, like Linux with their “tab is 8 spaces when deciding whether a line is too long” thing.
Tabs are self-consistent, which is usually all that matters. If for some reason you need greater control (E.G. ASCII-art spanning multiple indentation levels), then that should be documented instead.
Indent with tabs: To save time time adding/deleting characters or squinting at exact boundaries.
Better than both: Files are saved with spaces, but the code-editor is smart-enough to act as if the multiple spaces are a tab during editing or when setting a custom tab-width.
I always find that I’m clicking a bit too early in the “tab,” not selecting where I want, and then backspace is treating it like normal spaces again so things aren’t aligned.
I’d also love if it could render “tabs” of spaces at different widths, so that if one person wants to look at 2-space code, and another finds it more readable as 8-space, they have that option.
And anyway, all of this is only to help when a codebase is spaces when you want tabs, because I don’t see much benefit to saving out spaces if you’re going to go to all of this effort to emulate tab behavior.
Thank God I did not break my pinky.
For everything I write, that marginal speed improvement isn’t even close to worth the learning cost of 100% ditching the mouse. I’m bottlenecked on deciding what I want to express, almost never how quickly I can enter it into the computer. I just want it to not be a “clunky” experience, i.e. once I decide to select a block of code to move it somewhere, I want that to be painless.
For that, tabs give a nice large margin for error on the click targets, which means I can move faster.
The reason to use spaces is it keeps any alignment of multiline statements the same. You're not supposed to be able to adjust it because adjusting it means you can't do alignment, unless you use mixed tabs and spaces.
But as you say, you can still use spaces for alignment. I find it to be kinda meh because if the alignment has to change (like adding a longer variable name) you end up touching lines you didn’t otherwise have to in your diff.
Also, if we’re talking about editors doing nice things with tabs, I think you could make an editor look at the tab in and align things exactly how you want in rendering for a multiline statement.
I would be shocked if your editor didn't have a much easier option for you. For example, in Jetbrains tools I can just:
1. Double click on the first word to select it.
2. Shift click anywhere to the right of the final newline.
Or also: 1. Click anywhere at all on the line.
2. [Home] button to go to start of content
3. [Shift]+[End] to select to end of content.
That second one in particular resists any problems with shaky hands and squinty eyes.[Citation Needed, preferably with Actual Math]
I found one attempt to quantify the issue [0] which found differences which I don't think are very compelling, especially when you consider how insanely cheap disk space is compared to even a tiny time-savings in developer workflow.
Here, some napkin math just to put the orders-of-magnitude in perspective:
1. Assume the entire kernel repo is 1.5GB.
2. Assume your format shrinks it to 1.0GB.
3. Assume a $50 SSD with 1024 GB capacity.
4. ((1.5-1.0) / 1024) * $50 ≈ $0.025
5. Assume an *underpaid* developer at $20/hr.
6. $0.025/$20 * (60*60) = 4.5 seconds
So the disk-space benefits of reformatting to get the kernel-repo down by a (huge) -33% is wiped out if the (underpaid) developer ever spends more than +5 seconds dealing with the formatting.[0] https://www.reddit.com/r/C_Programming/comments/auv5mg/file_...
Also at work we all use 4 spaces for indentation. Nobody wants a smaller number. We have huge screens, we're not in the 80s anymore
I would like to fit more text on my modern screens, and 4-space indentation and 120+ char lines limit how many files I can view side-by-side while maintaining legibility.
I still try to keep my lines around 80 chars long in 2024 for this reason.
Unless you have huge peripheral vision too, your screen is likely not the issue.
Huge screens are useful to show more content e.g. multiple buffers, ancillary information (like docs).
I’m not interested in wasting that additional work space in emptiness because you can’t keep to reasonable line lengths.
I use 4 spaces and short lines because it naturally limits how far I can drift rightwards before I need to refactor or rethink what I’m doing.
Thought I’ll readily confess that rust can be mildly annoying there: lots of things require additional blocks, and the affordances of function scope around split borrows can make extracting utility functions awkward.
Wouldn't that be possible even when (exclusively) using tabs for indenting? Admittedly it's been a hot minute since I last programmed and I'm far more familiar with borland's ide with turbo c++ than VS code (so I don't know what the latest formatting works/features are), but if you only indent with tabs isn't that good enough? Also, do multiple spaces not take much more time to type out? Or have you remapped your tab key to multiple spaces?
https://www.reddit.com/r/programming/comments/p1j1c/tabs_vs_...
The last time I mentioned this on mastodon someone responded what I consider the best possible way to answer to all of this:
> The optimal tab width is e. It's wide enough to be obvious, but not so wide as to reduce line length too much. Plus, it discourages the use of tabs for alignment by being the number most difficult to approximate with any integer number of spaces.
... and when I say "best possible way" I mean "trolls everyone while making a solid point"
The tabs were then added back which I also think was the right thing to do.