Four Column ASCII (2017)
garbagecollected.org
garbagecollected.org
When you were punching a paper tape, you couldn't "delete" a character, unless you wanted to risk cutting and splicing the tape, and good luck with not jamming the reader.
But by convention, a RUBOUT character (with all holes punched) would be ignored. So if you made a mistake you could backspace the tape and punch RUBOUT.
RUBOUT had another useful purpose. It's well known that the DOS/Windows line ending convention (CR/LF: Carriage Return/Line Feed) comes from the days of Teletypes, but what is less known is that we didn't actually just punch CR and LF. Sometimes on a Teletype that wasn't perfectly maintained, the CR and LF would not give enough time for the carriage to settle on column 1.
So we always punched RETURN, LINE FEED, and RUBOUT. The RUBOUT would be ignored, and it added a little extra time for the carriage to settle into place and not get a blurry character in the first column.
Here's a nice picture (although the discussion isn't completely on the mark):
https://www.reddit.com/r/MechanicalKeyboards/comments/2v3k0p...
Many Compose key sequences are based on this.
So ASCII, in a way, may well be the first -or one of the first- variable-length codeset.
Note that overstriking upper-case letters to add diacritical marks does not work for ASCII (see more below).
Also, overstriking was how one typed diacritical marks back in the days of mechanical typewriters. Spanish used to (and may still? I forget) permit upper-case letters to not carry diacritical marks precisely because it was difficult or costly to get typewriters and printers to print such letters. Adding diacritical marks to upper-case letters requires fonts designed for that, and clearly an overstrike sequence cannot work unless the typewriter/printer holds printing the character until it knows whether the next character is BS.
It entertains me how many of these issues were addressed long before we had computers.
(Aside; could I argue morse was a variable-length encoding?)
Ah, this must be what you meant, yes?
—•—— • •••
I think some fonts are not great at rendering Morse code :)
There may be a dah dit dah dah dah, but I don't know it either. But then again I don't know the less common punctuation, Cyrillic Morse, and I think heard once there is even Kanji Morse.
...-.-
https://en.wikipedia.org/wiki/C0_and_C1_control_codes#Field_...
ASCII 28 for File Separator
ASCII 29 for Group Separator
ASCII 30 for Record Separator
ASCII 31 for Unit SeparatorBesides, it's arguably a GOOD thing that the delimiter in CSV (the comma) is such a common character, because it forces all parsers to properly support escaping and quoting. If it used some sort of unique character that was almost never used except in CSV, then 99.9% it would work without correct escaping, but the 0.01% when someone entered "you should use the character X to separate fields" as a column it would fail, and those cases are much more likely to slip through the cracks.
Also: the whole point of CSV is to be human readable. There's no obvious way to render "Record separator" on screen, and the format would essentially become a binary format.
Our business has parsed and emitted significant volumes of data have never encountered one of these code points in the wild.
My understanding of why not to use these code points is that they don't have printable glyphs and thus don't easily lend themselves to ad hoc data exploration tools.
It's my go-to ascii table when I don't have the `ascii` command-line tool.
In retrospect, maybe I did know of the existence of a command called `ascii`. It's the reason why I needed to specify the section to `man`, since it would otherwise take me to ascii(1).
It’s still located there on many international keyboard layouts:
Interestingly, shift-0 would logically be a space character, and indeed on the JIS layout there is nothing printed above 0 (actually pressing shift-0 doesn't give you a space on modern OSes, it just gives you a 0, but you could map it to a space if you were so inclined).
* except that ¥ is where \ would be in ASCII, because Japanese charsets replaced \ with ¥. On the JIS layout, the \ label is next to _ (0x5f), which means it's filling the empty spot of 0x7f, which is otherwise a nonprintable DEL.
It seems to me that it's only the elements in the third column that can be used to generate elements of the first column. This is consistent with the description in the article, and still explains why ESC is represented by ^[.
> Pressing CTRL simply sets all bits but the last 5 to zero in the character that you typed. You can imagine it as a bitwise AND.
If that were the case then CTRL-(any column) would result in the first column.
Edit: I suppose at this point terminal emulators are just hardcoding in CTRL modifiers rather than the truly emulating what a hardware terminal would have done.
> Pressing CTRL simply sets all bits but the last 5 to zero in the character that you typed. You can imagine it as a bitwise AND.
But it can't be quite that simple, or Ctrl-[, Ctrl-;, and Ctrl-{ would all have the same effect. There must (as you said, and contrary to the article) be additional logic that only zeroes the leading bits if they are "10".
As the sibling says, it’s probably the case that terminal emulators have this historic behavior for specific combinations, and not for the more general mechanism of merely setting the high bits of whatever symbol was generated. It’s probably a bit fallacious to test things in an emulator... perhaps the article is correct for the original systems?
Another special case: On most keyboards (I suppose it's the terminal emulator that implements it), CTRL-SPACE emits the NUL character. (I find that useful because I use C-@ as the prefix key in tmux.)
Some did a subtraction of the code and others did a bitwise AND.
Source: https://en.wikipedia.org/wiki/Control_character#How_control_...