Given this is fixed in RFCs, it is unlikely to ever change
None of that stopped UNIX from getting away from it, and the internet is basically built on UNIX. What transformations the C FILE stream implementation makes to linefeeds has nothing at all to do with TELNET, SMTP, FTP, or HTTP. These are very separate areas of concern. In Windows if you're writing socket code you have to explicitly write the \r and the \n, just like on any other platform.
Apparently Mike Muuss compared the two implementations, and his recommendation to go with Bill Joy's code was the decisive factor.
[1] https://www.youtube.com/watch?v=ds77e3aO9nA - I think that's the one
But to clarify on the salient part, I was trying to say that UNIX diverged (successfully) from what came before it, and the CRLF behaviour was from the teletype era that the early internet started in. My point about UNIX being a building block of the (modern) internet was more about pointing out that if UNIX could diverge from that and be a key component of the internet, it's absurd to think that Windows could not.
I can definitely understand how you got where you did from what I said, though.
Though this is more of an artifact of terrible specs and the IETF's silly love affair with "free" text formats. They actually take delight in showing off how crazy the encoding can be.
* http://homepage.ntlworld.com./jonathan.deboynepollard/FGA/qm...
file = open(path, "w", newline="\r\n")Microsoft must have thought to themselves in the early 90's: "Gee, we should support unicode... that means we need wider characters, right? Two bytes ought to be enough. Ship it!" Then when it became clear that two bytes wasn't enough, they didn't want to replace all the types for those win32 API calls, so they said "Well, let's just pretend that the two-byte encoding was actually UTF-16 all along", which is both variable width (which has all the problems UTF-8 has with invalid encodings), and it's wasteful in size (all ASCII text is stored with two bytes per character). It's literally the worst of both worlds.
Except it was Unicode which was a 16-bit code back then which got changed into 21 bits with Unicode 2 which also morphed UCS-2 into UTF-16. This happened in 1996, when Windows NT already existed. Blaming Microsoft (and Sun, and Netscape, ...) to follow a standard is a bit out of place. Heck, while UTF-8 everywhere would be nice, things are still messy and various different encodings are rampant. At least UTF-16 is immediately recognisable, as opposed to all the legacy encodings.
But at least in Unicode 2.0 in 1996, the definitions they finally gave were UTF-7 and UTF-8. UTF-16 didn't become standard until 3.0 in 1999, when it became apparent that they couldn't take the first 5 years back, and too much software was already written assuming 16-bit characters.
My point is that UTF-16 is a completely worst-of-both-worlds hack that nobody would implement in a vacuum, short of a need to maintain compatibility with a 2-byte character API. The only reason it exists is because people (the Consortium included) truly for a brief period thought a 2-byte character set would actually be enough, and then had to figure out a way to preserve compatibility later.
Meanwhile the UNIX world thankfully seems to have gone straight from ASCII to UTF-8 without any awkward dead-end in the middle [1], because if you need to maintain backward compatibility with anything, it's a lot more useful (and space efficient) for that something to be ASCII.
[1] Sure, there exists wide-character versions of all the posix standard APIs, but in practice all the other API's you're likely to use in UNIX land have converged around wrapping the 8-bit character versions.
That would really have been too much skeuomorphism in our text file format.
What is RUBOUT? It's a character with all 1 bits. On a paper tape, a punched hole represents 1 and lack of a punch is 0. So a RUBOUT character has all the holes punched out.
And by convention, RUBOUT is ignored when a tape is read.
If you made a mistake punching a tape, you would type BACKSPACE RUBOUT. BACKSPACE wasn't a character itself; it would physically back up the tape by one character position. RUBOUT then punched out all the holes, in effect erasing that character.
To go to the next line, you would punch RETURN to return the print carriage to the beginning of the line, and LINE FEED to feed the paper to the next line. But unless you were daring, you'd follow these with a RUBOUT to give the machine a little more time for all this mechanical movement.
While at it, I also wish if Linux distros could get done with case sensitiveness in filenames. Is there any advantage to allowing both Filename.ext and filename.ext? The only time I seen this is when a malware is trying to stay hidden. Also some naming conventions would be nice, when you can have any character including '/' and '.' as a filename, it gets annoying to deal with them in a cli.
Sorting by ascii order put capitalized files to the top of the list. Exactly where you should look for important stuff like README and Makefile.
Case preservation without case sensitivity is just dumb. Why bother, because it's pretty? Also, what's up with spaces in filenames? And furthermore, what are all these kids doing on my lawn?
What's_up_with_spaces_at_all? Let's_just_join_sentences_with_underscores_for_clarity.
Such an utterly minuscule amount though (especially in the modern times when storage and bandwidth are cheap and ubiquitous), I think you typing that sentence has wasted more storage and bandwidth.
I think a better reason is to avoid issues when working with code files across systems and version control. This actually causes major annoyance.
Sent a bazillion and one times all day, every day. I'm sure it adds up.
My point is that there's no shortage of digital storage space and bandwidth is only getting better. And if CRLF does cause a shortage of storage and bandwidth, I'm sure MS would address it.
The only widely used file system without case sensitivity is HFS[1]. NTFS to its credit is case sensitive, but Win32 is not, for legacy reasons (same reason as why MAX_PATH still plagues Windows).
[1]: https://en.wikipedia.org/wiki/Comparison_of_file_systems#Fea...
[1] https://helpx.adobe.com/creative-suite/kb/error-case-sensiti...
Yes, Unicode is messy and could have been better designed (it was designed so that there is an easy conversion path for any pre-existing encoding – thus concerned itself more with making it easy to convert content in legacy encodings to Unicode, instead of making it easy to implement applications in a way that they support Unicode), but it's still orders of magnitude better than anything that came before it when it comes to representing text in general. And it's mostly complicated because languages and scripts are complicated.
Yes, cat comes from simpler times, but if cat cannot be changed for compat reasons, then it should no longer be used to concatenate text files. At least if the result is somehow important. Text and binary data are simply two very different things and both need to be processed accordingly. Sure, there are a bunch of other operations that are immediately recognisable as not making sense on text at all and superficially concatenating files is not one of them, but in my eyes that's a bit shortsighted. You can safely use methods for binary data on text iff you know exactly what your text contains and that the operation is safe. Otherwise you may mangle things.