Imagine they automatically replace the "fi" with the styled "fi" behind your back and breaks text searches and normal text files.
You can be snobby all you want, as long as you initialed it and know you don't break programs' assumptions. No one says you can't have nice shit with LaTEX and Pagemaker when you need it.
Here people who just open the file and save it, or copy the text and paste it, don't mean it, yet it happens without them ever knowing.
As a reader, what is the difference between between fi and fi? Or between a and а?
If the search engine is distinguishing those because of Unicode characters, it has failed completely.
The longer I live, the bigger fan of simplification and standardization I become. Like with dates and times. Timezones and DST are a fucking nightmare, because politics. And then doubly so, because people enjoy themselves with their favourite regional writing formats. At this point I'm all for enforcing ISO 8601 on every communication involving dates and times.
Some people object that this is turning humans into machines, etc. So be it. Nature isn't perfect, and clear communications doesn't come to us naturally. Yet it is absolutely vital in a technological society.
Looking at the world as only “what will my Emacs store on the drive when I paste some text into it” is ridiculous. As computers become more advanced—even mobile phones are super computers now¹—we should use that computer power to make the technology work for us, not bend over to some 1980s concepts of computing and standards. It may be more difficult for you to build such a system, I understand, but as a used, I really don’t give a damn.
¹: https://www.macrumors.com/2017/09/13/a11-bionic-chip-geekben...
Yeah, Because we already changed " to 69 behind their back, let's double down on correcting the 69 so it processes like "!
Let's engineer our dumb, close-minded O(n) string search and CSV parser to do some AI image recognition shit in O(2^n) to figure out when the fuck the "used" had the ` involuntarily changed to ' because MacOS or Wordpress decided it looks better.
That's brilliant engineering right there.
Satire aside, you realize that there is a cost of security and maintenance to over-engineer shit, it's just not just convenience, right?
Your specific bug with CSV is a developer bug. The place where you copied the incorrect CSV format should not have had such substitution enabled. On the Mac, developer can specify for each text view and text field what substitutions are allowed by default. Likewise in web, it is possible to specify which substitutions should be allowed for text areas.
So instead of blaming incompetent developers for their incorrect use of system features, thus ruining some very narrow cases, let's hold back any kind of text input and processing advancements, because you are unable to input some CSV properly.
That's brilliant engineering right there.
When my mother types on her computer, she just wants things to work. When she searches for something, she doesn't care if she typed the wrong incorrect Unicode character. That's the cases that need to be solved for users.
And it is only O(2ⁿ) algorithm if you naively look at text as an array of bytes. Time to, perhaps, broaden some horizons.
Or she cares more when the shit she copy-and-pasted from some random website or your note to her to do something doesn't work because the site or the copy-paste process meddled with it?
PS: The joke is on the O(2^n) with AI Tensorflow, not on the "used." The "used" thing was 101% serious.
'n'
can either be an abbreviation of "and", or a single-quoted letter "n".as an abbreviation of "and", it should be rendered as [right single quote]n[right single quote]
’n’
and as a single-quoted letter "n", it should be rendered as [left single quote]n[right single quote]
‘n’
What I mean is, for example, this: when I see " in a comment, I know the browser is actually storing the " character somewhere in memory. I know that when I send this comment form, HN will receive " character. If I copy and paste the comment into my Emacs, and save it, I know my hard drive now stores the " character.
When external format gets completely disconnected from internal format, understanding anything about what happens with the data gets much more difficult.
There are (small or big) differences among languages and it is not always obvious to detect ,if the quote should be converted at all, if it's a left one, or a right one.
As a reference, this is a Python Script trying to do the conversion in Scribus:
https://wiki.scribus.net/canvas/Convert_Typewriter_Quotes_to...
There is also an article by the author of the script, explaining his work and pointing to at least one common case it cannot handle:
https://opensource.com/article/17/3/python-scribus-smart-quo...
Putting all this logic in a font might or not be something your really want...
But I 100%: this should be done at the font level... and hard replacing characters is not a good solution.