Elastic Tab Stops (2017)
nickgravgaard.com
nickgravgaard.com
If we're lifting abstraction one level up - from dumb plaintext to delimited records - maybe it's time to revive the ASCII control characters that were specifically meant for this? I'm talking about ASCII 28 to ASCII 31 - file, group, unit and record separators.
i think you should be able to nest multiple lines within such a 'table cell' by enclosing it between form feed and vertical tab, which are both conventionally considered whitespace, and are therefore allowed already in the grammars of conventional free-form languages like c, java, js, and scheme. nesting tables within table cells permits almost infinitely flexible layouts, as we all remember from the days before css; and, although the layout algorithm is not as trivial as emulating a teletype, it's fairly simple, linear time, and very fast (as long as you don't try to add word wrap or something)
(you probably also want rowspan and colspan; these should be indicated by putting carriage return and backspace, respectively, in the cells spanned into, like \^ in tbl)
a 'terminal emulator' using such a format would probably need an alternative to character-cell cursor addressing; probably the best alternative is to have multiple cursors, with one escape sequence to move cursor X to the active cursor (forgetting its previous position) and another escape sequence to exchange the active cursor with cursor X
there's no need for escape sequences to look like line noise either
further details are in nestable.md in
git clone http://canonical.org/~kragen/sw/pavnotes2.gitThat's exactly what's happening in currently rudimentary implementations like in Sublime Text, where width can be smaller if needed for alignment
The "spec" also doesn't seem to mandate a minimum
int someDemoCode( int fred,
int wilma)
i think it would be better to represent that, where actually desired, by 00000000: 696e 7420 736f 6d65 4465 6d6f 436f 6465 int someDemoCode
00000010: 2820 0969 6e74 2066 7265 642c 0a09 696e ( .int fred,..in
00000020: 7420 7769 6c6d 6129 0a t wilma). .
(note the explicit space following the left paren) and render 00000030: 696e 7420 736f 6d65 4465 6d6f 436f 6465 int someDemoCode
00000040: 2809 696e 7420 6672 6564 2c0a 0969 6e74 (.int fred,..int
00000050: 2077 696c 6d61 290a wilma).
as the more acceptable int someDemoCode(int fred,
int wilma)
which i think differs from gravgaard's intended semanticsBut then in your design how would you Tab to actually add whitespace, not just vertically align? Would then 2nd tab have a width of 1 space?
https://en.wikipedia.org/wiki/Tab_key
> The word tab derives from the word tabulate, which means "to arrange data in a tabular, or table, form."
So, say, for TSV files (where you can't add spaces as that would affect the field values) you'd just have an editor with this feature let the max-width columns be glued, correct?
And speaking of wiki, the key part of tab is this (the etymology is just a historical artifact that should have no effect on what the best design is):
> Pressing the tab key would advance the carriage to the next tabulator stop.
The tabulation stops is another great missing feature from the MS Words of the world
if you have a header line, and your field names do get trimmed on input (or discarded), you can insert the padding spaces there
for tabular data display, escape sequences to insert vertical rules between columns are also desirable
the tab-stop mechanism is just a 19th-century limitation in the mechanical typewriter technology of the time. we can design new systems today that do what we think is most useful; from my point of view, arranging information in a table is more useful than advancing to the next tabstop
here's some output i got from vmstat 5 today
procs -----------memory---------- ---swap-- -----io---- -system-- ------cpu-----
r b swpd free buff cache si so bi bo in cs us sy id wa st
0 0 1557232 132264 121644 1632328 1 4 34 48 17 57 12 5 83 0 0
0 0 1557232 132004 121652 1632368 0 0 0 9 561 971 1 1 98 0 0
1 0 1560816 233716 119836 1534920 0 717 0 4020 1286 2227 13 5 82 0 0
0 0 1560816 234220 119844 1534920 0 0 0 10 863 1919 3 2 95 0 0
ughas it happens, arranging text in tables also gives you the ability to do a flexible visual layout that accommodates proportional fonts of multiple sizes, so you can also stop using ugly typewriter fonts
we can stop imitating the typewriter's flaws now, we've had pixels for 60 years already. the question is how to make pixels as easy to use as a typewriter
(not sure how escape sequences help in the editor, that would still be extra manual chars inserted into each column?)
I think a better approach would be for the editor to zero tabs in semantic locations like function(arguments)
The tabstops aren't mechanical limitations in apps, the allow you to manually adjust the offset, which works great, eg, for comments (set all comments at a tabstop at column 80 even for not-contiguous blocks of code, where elastic tabstops wouldn't work, although proper design would also need semantic awareness not to conflict)
you could have your editor insert the trailing space automatically (perhaps with exceptions for cases like function(<tab>arguments)), then manually delete the space in the cases where you don't want whitespace between table columns; this might even be a good default for things like editing c and js, which is kind of annoying to do without editor support for auto-indentation
you could very reasonably have a single escape sequence that sets a minimum padding for all the columns present on the current line; prepending it to the first line of a tsv file on output, or appending it to the last, might cover the tsv-editing case relatively well. or you could nest the tsv file inside a form-feed/vertical-tab pair, and define the escape sequence to apply to everything inside that cell
my concern is to make the format able to express general layouts, not to make it maximally compressed
manually adjusting tabstops is only an adequate replacement for a general string representation for nested tabular text layouts in a few niche special cases like a single developer looking at their own code. it doesn't help vmstat or ls to produce more readable output, enable a program to use multiple font sizes in its output without worrying about font metrics or html reflow speed, make tabular data easy to edit, replace the windowing system, or even enable two developers editing the same plaintext code on different computers to format it so the other person can read it easily
vanilla elastic tabstops are i think not very well suited to keeping comment blocks on the right aligned across multiple indentation levels of code; my proposal of nested table cells can do that in the special case where all your lines are the same height. alternatively, even with gravgaard's vanilla proposal, you can keep your comment blocks aligned, as long as you put them to the left of your code and not the right
btw i saw that someone downvoted your comment, so i upvoted it; hopefully that keeps it from falling back into the grey. you're bringing up some important considerations even if not all your points are well thought out ;)
That's just as bad, manual fiddling is precisely what autoformatting algorithms are supposed to solve. Especially in the more painful cases of tables. Also, inserting many spaces in large tables is a rather costly operation, I have to use debouncing in the current elastic tabstops plugins that align with spaces to avoid that cost on every keypress. Also, you're asking to differentiate between the two invisibles (tabs and spaces), that would require making one of them visible, but then in many cases that visibility would be legal and thus just visual noise
> you could very reasonably have a single escape sequence that sets a minimum padding for all the columns present on the current line;
You can't do that since in many languages these sequences would be invalid input
> my concern is to make the format able to express general layouts, not to make it maximally compressed
Yet your format can't do that without a lot of manual fiddling
> manually adjusting tabstops is only an adequate replacement for a general string representation for nested tabular text layouts in a few niche special cases like a single developer looking at their own code
This also works in such niche special case as non-code editing
if you only support one of the two options, it's true that you don't need that user-interface action, but that's not because it requires less manual fiddling to achieve the same result; it's because it can only achieve the result that the more expressive format could achieve without manual fiddling, while the other option becomes impossible instead of fiddly
and that's true regardless of whether the fiddling-free option has the extra spacing or lacks it
you seem to have been responding to an earlier edition of my comment, by the way, because your response no longer makes sense in context because i already clarified some things you misunderstood; i mention this partly in case other people are puzzled as to what you are talking about
i didn't realize you'd implemented elastic tabstops yourself! which editors did you write it for?
i don't agree that relayout of a large table with new column widths is an inherently expensive operation (in the absence of word wrap) but i can easily believe that there are software environments that make it so
The alignment logic should be separated from the appearance: use some codepoint to mark each cell of a column, calculate the minimum x from all cells of the column, and then - when displaying text - apply eventual font hinting to round up (right) this x…
i wasn't suggesting manually positioning text pixel by pixel, but rather taking advantage of the potentials unlocked by displaying text with pixels rather than a teletype
This won't do much to sway these programmers
We space users all use the tab key. We just make sure our editor replaces them with spaces as we type because of decades of bitter experience with how mixed tabs and spaces behave over time.
The more interesting question is why tab users haven't noticed that we use mandatory auto-code-formatting these days, to make sure all your tabs are belong to us.
;-P
I'm guessing they're using an equivalent of `M-x tabify` running automatically after opening a file.
Excerpt from documentation on `tabify` in Emacs:
Convert multiple spaces in region to tabs when possible.
A group of spaces is partially replaced by tabs
when this can be done without changing the column they end at.
If called interactively with prefix ARG, convert for the entire
buffer.
Personally, I'm a "space user", and I only have `delete-trailing-whitespace` run on save (`before-save-hook`); at work, we have a pre-commit hook that untabifies committed changes (and then another to run language-specific autoformatter).Not all. I've seen people hit spaces repeatedly…
a = 1
ab = 2
abc = 3
becomes a = 1
change = 2
abc = 3
Nobody ever said a word about it.In Elixir projects we've been using the language code formatter for years and it forces a coding style like mine by default.
Honestly, I use spaces because virtually all tooling and coding conventions that exist prefer spaces over tabs, but I don't understand how it came that spaces won over tabs in the programming world. So much energy wasted on counting spaces in text documents, it's absurd.
Who on earth is counting spaces?
If I'm in Python I can press tab multiple times to cycle through indentation levels or do it manually with space and the deletion keys.
We use spaces for indentation. Do not use tabs in your code. You should set your editor to emit spaces when you hit the tab key."
https://google.github.io/styleguide/cppguide.html#Spaces_vs....
Although most coding guidelines seem to have converged towards spaces over tabs (see [1], [2],[3], your source, and many more), it also causes accessibility issues, especially for visually impaired people.
Some people require larger fonts, making space-indented lines go offscreen quicker. Some people require braille displays, which have a limited line length. A tab takes a single character on these, vs multiple (usually 4) for space-indented code. Some people require bigger indents to be able to process them better.
Not even considering disabilities, some people prefer 2 spaces, some 4, and some whatever pleases them. Working with tabs allows everyone to work on the same codebase with their own preference. By definition this is a subjective matter, and trying to enforce a preference across a whole codebase or programming language is needlessly opinionated.
Plus, the 2 first example use case in the article can't be solved either with spaces or standard tab stops.
---
[1] https://peps.python.org/pep-0008/#tabs-or-spaces
> Spaces are the preferred indentation method
[2] https://learn.microsoft.com/en-us/dotnet/csharp/fundamentals...
> Use four spaces for indentation. Don't use tabs
[3] https://www.oracle.com/technetwork/java/codeconventions-1500...
> Four spaces should be used as the unit of indentation. The exact construction of the indentation > (spaces vs. tabs) is unspecified. Tabs must be set exactly every 8 spaces (not 4).
While leaving tabs vs. spaces "unspecified", specifying stabstops of 8 spaces actively discourages tabs
>Working with tabs allows everyone to work on the same codebase with their own preference.
This is where the problem starts. As someone who has to integrate code from a dozen devs, it seems that everyone has a sligthly different usage of tabs (TFA adding one more). If we can ever set everyone on the same page with tabs, I am fine with it. It's just that in production it always seems to cause problems. Spaces are easier to specify.
For the case of disability support, I'd risk to guess that different usages of tabs would end up derailing accessibility software just as much as it derails integration. But I'd prefer deferring this opinion to someone who has experience with it.
Do you want to say: Google bad, Google good or what?
- do not indent code that way
- do not align comments on the side
Before
int someDemoCode( int fred,
int wilma)
×(); /* comment */
print("hello again!\n"); /* comment */
makeThisFunctionNameShorte(); /* comment */
After int someDemoCode(
int fred,
int wilma
)
×(); /* nobody */
print("hello again!\n"); /* needs */
makeThisFunctionNameShorte(); /* aligned comments */
This kind of control over indentation and alignment borders on OCD. You should be using a code formatter anyway, so indentation is auto-fixed. int animateIndenting(
int startingIndent, /// Starting indent, in spaces
int endingIndent, /// Ending indent, in spaces
int animationTimeMs /// Animation time in milliseconds
) {
...
}
is tempting. It's an interesting idea. Whether it's a good idea. It's hard to be completely objective when, for a brief time in the 80s, I spent one third of all my coding time maintaining elaborately baroquely indented Pascal comments. (Not my coding standard. It was fashionable at the time).That the formatter doesn't replace tabs with space... quaintly optimistic. A great shame, since requiring that people use editors with this feature for all eternity, or have ridiculously random indentation of comments is a non-starter.
GP’s example is just finicky window dressing based on indulging a personal aesthetic based on what the author things “looks good” on a particular day, with a particular font on a particular screen.
What we should be striving for in code style is the lowest common denominator most bare bones strucure that can be read, understood and modified with the most rudimentary of tools.
EDIT > What happened to "code is primarily to be read, occasionally to be executed"?
Exactly this. Code should be readable for all. It goes through all sorts of other actors and tooling, not just your painstakingly configured IDE.
>> What we should be striving for in code style is the lowest common denominator most bare bones strucure that can be read, understood and modified with the most rudimentary of tools.
> as if, when designing a building, everyone involved - architects, plumbers, electricians, HVAC engineers …
All the above have to abide by building codes, use standardised parts and tools and most have to conform to certified standards which their work is then also validated against.
We have many, many coding standards most of which discourage tabs and fancy formatting but developers get special exemption because “ma creativiteh”
> plaintext code directly is an idea that needs to die.
Yes let’s throw out some things that work perfectly well to accommodate a few malcontent crybabies.
At the end of the day our laws and most of our written communication depends heavily on plaintext and that’s not gona change any time soon just because a few people don’t want to learn how to use it properly.
What happened to "code is primarily to be read, occasionally to be executed"? You're suggesting we should be optimizing for the wrong thing.
To be honest, this is just another unsolvable issue that's a direct result of us insisting on keeping the code as plaintext and insisting we work directly on it and only it, the single source of truth, using bare-bones tools.
It's really as dumb as if, when designing a building, everyone involved - architects, plumbers, electricians, HVAC engineers, decorators, landscapers and others - were forced to use the same, single, canonical, flat technical drawing. You can imagine they'd be wasting much time on endless holy wars on the right drawing scale, right colors, line thickness, where to put the legend, how to make annotations, etc. - where the real problem is that a single flat piece of paper (or digital equivalent) can't possibly accommodate all their needs at the same time.
We're like that, and here we're discussing whether to align labels on the drawing to the left or to the right, in hopes it'll somehow make an impenetrably dense drawing an iota more readable. We're going as far as to play with dark monadic magick, inventing convoluted drawing techniques in order to somehow squeeze layout of cables, water pipes and air vents in a way that makes decorators shut up about not being able to visualize the building in their heads, because the plan is too dense or too non-local or whatnot. And we're so proud that we automated the ability to xerox the plan for different teams to add more squiggles independently, and then merge them into a new unified plan the next day, and even have an audit trail by putting SHAs on everything.
Seriously, our whole approach and tooling ecosystem is tripling down on idiocy. Editing the same, single source of truth, plaintext code directly is an idea that needs to die.
I use Obsidian, because it's a cool notepad software that stores everything as plain text. This means at any point in time, if Obsidian stops being developed (and supported by up-to-date OSes), or there's a new better software, or stops being free to use - I can just stop using it, and I still have my notes. This is just one of many strengths of the plain text. Compatibility is another.
I know, I'm not actually editing those plain text files directly. I can, however.
Plain text as representation to read, write and edit is also fine - many things are best expressed this way, and text lends itself to be edited efficiently by advanced tools (e.g. vim, Emacs, or sed, awk, etc.)
The issue is not with either of them alone. The problem seems from using the same plaintext documents for both storage and representation simultaneously.
The solution is to have a storage format optimized for storage - which could just as well stay plaintext - but all reading, writing and editing to be done through tools that render you a representation best suited for what you need at the moment. This could be plaintext prose now, plaintext code an hour from now, and a pile of editable, graphical state machine diagrams tomorrow.
Plaintext is for interchange not storage. You can optimise for storage by putting .gz on the end of it.
> writing and editing to be done through tools that render you a representation best suited for what you need at the moment.
XML? XSLT?
This has been tried and it’s easily done if you want nobody to like you.
The crux of the issue is interchange. Neither of the things you have so vociferously decried. As usual, somewhere in the middle.
The lowest common denominator is plain text and that’s what works best en masse for all the participants and tooling.
This is a huge logical error. You just proposed to use the lowest common denominator, but then your response to the analogy is that other people are constrained by rules. Yes, they are constrained by rules, that are very complex, and even then, the constraint ensures a minimal standard, it doesn't block you from creativity. Fire prevention rules will define how wide a corridor has to be, but it won't forbid you to build a wider corridor or to paint its walls.
BTW editing your post to answer a post below you wasn't the smartest thing to do…
The code formatters aren't good enough (and also elastic tabstops work realtime)
What matters much more are idioms, matching interfaces and descriptive names.
Once you get over that as well, it's also quite liberating to not be bound by rules. This does go into unlawful-chaotic territory for some people though
,