What actual purpose do accent characters in ISO-8859-1 and Windows 1252 serve?
retrocomputing.stackexchange.com
retrocomputing.stackexchange.com
Ultimately, many "character sets" combined printable characters, cursor/head control, and record boundaries into one serialized byte stream that could be used by terminals, printers, programs for all sorts of purposes.
Any such solution will always be broken, proof is left as an exercise for the reader. Hint: use Ogden's lemma.
So if you decide not to support "strings", just "text", then those separators Just Work.
But if I was writing a similar standard today and someone was pressuring me that I really had to make some kind of mechanism for nesting, I'd use something that transforms the field to remove the need for escaping, in a big and obvious way. Probably base64. That prevents most of the implementation issues you see with CSV.
CSV needs escaping, so don't do CSV.
This difference carries over to C, where NULL got the job done for string termination pretty darn well, even if there were strong critiques to be weighed against it.
From a purely linguistic standpoint, the choice is sound. Obviously it's objectively a disaster but for unrelated reasons.
They can, only some functions stop at the first zero:
char s[] = "gogo\x00gogo";
printf( "%d\n", (int)sizeof( s ) );Consider the following program:
#include <stdio.h>
#include <string.h>
int main()
{
const char * const s = "gogo\x00gogo";
const char t[] = "gogo\x00gogo";
printf("%d,%d,%d,%d\n",
(int)sizeof s,
(int)strlen(s), // DO NOT EVER USE THIS FUNCTION
// YOU WILL BE FIRED IF YOU DO
(int)sizeof t,
(int)strlen(t) // DO NOT EVER USE THIS FUNCTION
// YOU WILL BE FIRED IF YOU DO
);
}
> 8,4,10,4> They can, only some functions stop at the first zero:
> char s[] = "gogo\x00gogo";
And in what human language is that `\x00` a valid character? What rune/glyph/symbol represents it?
Nothing, you say? Well, looks like a good choice to me.
I think a big reason they never took off is also that there's no visual representation of them, or input method for them.
I'd be happy to use them instead of CSV if I could edit records in Notepad or TextEdit the way I can with commas and quoted strings. See the separators, type them.
But of course once you do that, somebody now wants to insert a list of five values separated by unit or record separators into a field. They'd just be common ASCII characters along with tab, CR, LF, etc, and need to be escaped the same.
There's no escaping escaping...
> There's no escaping escaping...
This is the truth.
And in practice I do not expect CSV with embedded nulls to work properly, so there's already precedent to reject certain characters entirely in a CSV-like format.
CSV is more broken than you imagine. Try opening an American CSV file with Excel in Germany or France. You’ll discover that it doesn’t parse, and the reason is that the ‘C’ in CSV apparently stands for “could be a semicolon too” because comma is the decimal separator in these countries and therefore must be available for use inside values.
- Make sure to write your own parser! The format is so simple!
- Is the first line a header or not? The RFC says you should look at the MIME type to tell. You can't make this up.
- Is the line terminator "\r\n", "\n" or "\r"?
- What happens if the number of records for a line is incorrect? Well, that SHOULD'nt happen...
- Make sure to involve Excel, nothing ever went wrong feeding data to Excel.
- Are end-of-line spaces relevant or not? Should we trim them?
There is no defending of CSV, it's a broken format and it's broken beyond repair. The point is that the separator is immaterial. Replacing commas with something else gains us exactly nothing, and results in an identically broken format.
For one, I think you'd almost certainly use Record Separator, and \r and \n would be normal text characters inside of a field. And no trimming spaces.
And the behavior upon encountering corruption isn't really up to the spec, is it?
Headers might be more trouble than they're worth, but if they're in I'd say they need to be mandatory. One code path.
You'll find yourself with a much less broken format.
Always the seemingly simple solutions have most issues in real world... At some level still trying to edit stuff by hand might be where we have gone wrong for long time. Why don't we have any sensible agreed structured format for which tools would be on all and every device...
> ASCII separators are equally as broken. How do you represent a string containing the record separator in a record-separated file?
ASCII separators may not be perfect, but they're certainly less broken than CSV. With CSV, you have the pervasive problem of how to represent a common, everyday character in the file. With ASCII separators, you only have the problem of representing an unusual, untypeable character, which is much less troublesome.
It's the temptation to just concatenate or split based on commas, or to hack a broken implementation of quotes on top, that causes so many issues.
There are two main options for that. Make it so that poorly parsing with self-written code is still very likely to do the right thing, or make the parsing hard enough that people go get a library.
Printers were fare more common than hardware terminals for most of the initial history of computing.
https://en.wikipedia.org/wiki/Dead_key
I can't claim I know how they relate exactly to the accented characters being encoded in character sets but they seem to be at least historically an influence. Pressing a dead key which doesn't advance the cursor and then overwrite the basic character over it is certainly faster than using backspace (also cheaper if you think about character pricing).
That the ECMA specs only talk about using BACKSPACE is surprising. At least those OS I used only supported the dead key approach but of course that was decades after the specs were written.
If the expanded count doesn't match, a diacritic might be present.
[0] https://learn.microsoft.com/en-us/dotnet/api/system.text.nor...
To some extent the character set was still evolving, for example the Euro sign was not around until decades later and that would need to be bolted together with backspace characters or escape codes, maybe even downloaded characters, with the printer specific manual (Epson) studied at great length.
In the DOS era (and before with home micros that were programmed in BASIC) it was quite normal to compose things for the printer that you had no expectation of seeing on screen, not that anyone read much on screen (as everyone had vast piles of paper on their desk).
Until quite recently some POS systems were very much tied in to a very specific printer, at least these character sets were a step forward from hard coding a BASIC program to an exact make and model of printer.
https://en.wikipedia.org/wiki/Ditto_mark
Guessing maybe there’s some keyboards which were missing quote/prime characters but had umlauts?
I was coding forms for tabular data, doing silly things because the printed manual made that too tempting, so any excuse to use a character or escape code sequence was fair game.
This was in an era when handwriting was commonly used so the convention for 'ditto' could be freestyled by me to be middle dots, umlauts or anything CHR(). The convenience of the keyboard, in an era before the PC was not really there.
The asker completely ignores that asking questions about accent marks, like they themselves are doing in that very post, would be a lot more annoying without being able to write said accent marks.
You can use it for referring to the new Glőrbl character that I invented.
Otherwise you would have to resort to the quite clumsy "the thing on top of àìòèù" to introduce any discussion of the uses of grave accents to people who don't yet know what the phrase "grave accent" means instead of just writing "`", which would be akin to saying that we don't need to write the letter "A" by itself to talk about the letter "A" and could instead say things like "the letter at the beginning of the words Apple and Among". Like...yes...you can do that. But it's less good than the alternative of not doing that.
For example, one could type ë by entering ¨ then following with e. The ¨ would be displayed at the position where the combined character would be, while waiting for the second character to be entered. Once the second character is entered, the display would be updated with the correct combined character.
New? Some computers of the 1980s could already do this. At least the 16 bit home computers had bitmap drawn characters on the screen.
Edit: Looks like somebody doesn't believe that computers in the 80s had such a thing.
> On the Amiga, rendering text is similar to rendering lines and shapes. The Amiga graphics library provides text functions based around the RastPort structure, which makes it easy to intermix graphics and text.
> In order to render text, the Amiga needs to have a graphical representation for each symbol or text character. These individual images are known as glyphs. The Amiga gets each glyph from a font in the system font list. At present, the fonts in the system list contain a bitmap of a specific point size for all the characters and symbols of the font.
At least 40 years is not very new.
Otherwise HYVÄÄ YÖTÄ was HYV{{ Y|T{, which was only little miserable.
But if you changed the ROM into Swedish ROM, {a|b} become äaöbå, which was basically unreadable.
Printers are an obvious answer for a different question.
Edit: Its 1983 predecessor that added accented characters was made for the VT220, and it used a buffer of 8 bit character codes, no way to combine. (Some of those terminals could be sent a limited number of custom characters to store in RAM, but that's a very different mechanism as far as I can tell.)
Edit 2: I made an error before, the naked accents were not in the predecessor, they were even newer.
For that, the printer would also need to interpret the characters in the correct character set.
In 1985 "OS Support" meant picking up the phone and talking to the vendor directly.
The answer is that the character sets were designed with specific hardware in mind and often with backwards compatibility.
The accented characters were in use for typewriter and printer with a typewriter like interface or even hardware.
Someone in 1985 who has mostly printers. ASCII dates from the 70’s.
Yes, and ASCII did not have these characters.