And the funny thing is: it's not that complex.[edit: let me rephrase: it is, but it doesn't have to be. C family is a nightmare, Python et al are delightful]. Unless you don't know how it works---then it's the mystery Gordian knot, as you describe it.
The irony is that encoding is a worry precisely for those who try and stay away from it.
Don't shy away from encodings; embrace them. Then you will learn to love them.
(Another irony: with UTF8 gaining more and more mind share, encoding issues actually become harder to find and debug: they don't show up, and when they do, fewer and fewer people know how to deal with them. Everyone switching to UTF8 just hides the bugs, until it doesn't.)
Unless you do multiplatform development, then the language has a hard time saving you (and Python definitely does not)
> (Another irony: with UTF8 gaining more and more mind share, encoding issues actually become harder to find and debug: they don't show up, and when they do, fewer and fewer people know how to deal with them. Everyone switching to UTF8 just hides the bugs, until it doesn't.)
That's not true at all. A ton of byte sequences (and 13 standalone bytes) are outright illegal in UTF8, there's a fair amount of error handling in a validating UTF8 decoder[0], whereas there usually isn't any invalid byte (let alone byte sequence) in 8-bit codepages or character sets. When you decode random bytes in Windows-1256 or ISO-8859-9[1] and re-encode them as UTF8, the UTF8 encoder isn't the one at fault for your garbage output, as far as I could see it got perfectly valid unicode data.
People passing through unvalidated data (possibly assuming it's UTF8) isn't a problem with UTF8 either, by the way.
[0] whether that's used in strict mode or in replacement mode is a different concern and not a blemish on UTF8 itself
[1] which will always succeed, you can try it at home, just get a bunch of bytes from urandom and feed them to various decoders, chances are low that you'll generate anything the UTF8 decoder will accept, chances are also low that you'll generate anything an ISO-8859 character set will reject.
It's sad that a "modern" OS that had a mostly ground up rewrite (NT) after utf-8 was invented, doesn't have better support for it. I get in memory storage being utf-16, and I'd even accept modern OSes storing files in utf-16, but utf-8 is so elegantly backwards compatible with ascii, it's brain dead not to fix -everything- to work with it.
Perhaps this is the biggest difference between a closed OS like Windows and it's more open counterparts, if Windows were open then the community could have fixed this issue long ago.
> The ASCII standard has nothing to say about backspace overstriking
It seems impossible to get ANSI copies of old standards (even for money) but the ECMA (1973) printing says: 3.2 Diacritical Signs
(Positions: 2/2, 2/7, 2/12, 5/14, 6/0, 7/14)
In the 7-bit character set, some printing symbols may be
designed to permit their use for the composition of acce‐
nted letters when necessary for general interchange of
information. A sequence of three characters, comprising
a letter, BACKSPACE and one of these symbols, is needed
for this composition; the symbol is then regarded as a diacrit‐
ical sign. It should be noted that these symbols take
on their diacritical significance only when they precede or
follow the character BACKSPACE; for example, the symbol
corresponding to the code combination 2/7 normally has the
significance of APOSTROPHE, but becomes the diacritical
sign ACUTE ACCENT when preceded or followed by the character
BACKSPACE.
This is precisely the reason ASCII 1967 replaced ← with _ and ↑ with ^ (explicitly still “circumflex accent” in Unicode) and added ` (“grave accent”, likewise).ANSI made this optional in the 1986 revision, in §2.1.2 — “The use of BS for forming composite characters is not required.” — with a note that it would likely be removed from a future revision (but there never was another one).
> In a more common example, ASCII has nothing to say about how you move the cursor to the start of a new line.
It does; it just says something slightly unfortunate about code 0x0A. CR Carriage Return
A format effector which moves the active position
to the first character position *on the same line*.
LF Line Feed
A format effector which advances the active position
to the *same character position* of the next line.
[Italics added] But then it says: The Format Effectors are intended for equipment in
which horizontal and vertical movements are effected
separately. If equipment requires the action of
CARRIAGE RETURN to be comhined with a vertical movement,
the Format Effector for that vertical movement
may be used to effect the combined movement. For example,
if NEW LINE (symbol NL, equivalent to CR + LF)
is required, FE2 shall be used to represent it. This
substitution requires agreement between the sender and
the recipient of the data.
The use of these combined functions may be restricted
for international transmission on general switched telecommunication
networks (telegraph and telephone networks).
So CR LF will unambiguously get you the first position on the next line. The code for LF is allowed to be replaced by NL by “agreement”, but CR can't move to the next line.Windows NT was released in July 1993, but development started in 1989 (https://en.m.wikipedia.org/wiki/Windows_NT#Development)
Also, even after UTF-8 was accepted to be a good idea, it wasn't considered a good idea for in-memory usage; indexing UCS-2 encoded strings is way easier, and UTF-16 didn't exist yet. It arrived with Unicde 2 in July 1996 (https://en.m.wikipedia.org/wiki/UTF-16#History)
No. In python you have the absolutely retarded behaviour of a program running fine in the command line but crashing if you redirect the output to a file.
C is much better, because at least it doesn't pretend to support encodings. So you're forced to use a 3rd party library anyway.
C is better because it _doesnt_ provide some kind of interface into encoding. A byte is a byte is a byte. Whether that's some slice of a binary blob or the first byte in a multi-byte unicode string is completely irrelevant. The most infuriating thing I've found with higher level languages (Ruby/Python/etc.) is their abstractions on top of string encodings that fail to cover the hundreds of thousands of edge-cases (which, they rightfully shouldn't cover either), meaning that the only time you do end up having to wrangle encodings is when things go south and you have to fight the language's defaults just to get some stupid email address that most likely had a bit flip in transit somewhere to properly write to a CSV.
The bigger problem is when Excel tries to be helpful and reformats the data upon opening the file. I can't remember the intricacies, but it used to behave differently whether you opened it from a local path or directly from a browser. I seem to recall using an "incorrect" MIME type helped with that.
Excel CSVs default to the system codepage when reading and writing, which makes some sense given that the files are plain text. You can actually control what codepage is used when importing data but you have to use the import option.
I found many of the protocols to be both incredibly clever and incredibly annoying at the same time.
What is a fixed delimited file? A 'mainframe' file is generally fixed length or delimited.