This post is a classic on various name issues: http://www.kalzumeus.com/2010/06/17/falsehoods-programmers-b...
This post is a classic on various name issues: http://www.kalzumeus.com/2010/06/17/falsehoods-programmers-b...
Explanation:
echo "é" | iconv -t utf8 -f iso8859-15
We're Dutch, and the é is part of our language, and even part of the legacy character encoding standard everyone used before Unicode's widespread adoption. This is just a matter of code that works perfect as long as all characters are part of the ASCII set, but fails on the characters that don't conveniently match between UTF-8 and ISO-8859-15.I doubt these issues will go away within even, say, twenty years.
Here in CJK territory, using the wrong encoding makes the output so obviously broken [1] that mistakes are almost always caught before hitting production.
If you happen to be using Windows configured in a foreign language the first time you start Outlook, your inbox, sent mail, etc, folders are named according to that language, and will never change, and you'll have to live with non-standard names for the folders.
At least its teaches you how to configure folders manually in most email clients.
For anyone who doesn't know OSX, this translation happens on the UI level. Typing ls in a terminal gives you the real directory name.
Yes. In some languages those are actually not "markings" but denote proper letters, like in German ä,ö,ü and ß. But even if not, like in French, it can alter the meaning of words. E.g la != là. Therefore, for most Europeans and speakers of other languages that depend on more letters than ASCII provides, it is very annoying when that is not supported properly.
However, I have made the experience in a few cases that particularly Americans have a hard time understanding this. The remark about your wife not caring seems to be in this vein, too. Recently, I decided to convert our MySQL DB tables from latin1 to UTF8. (I wasn't even aware that we didn't have some form unicode, as our DB is only few years old, and I thought some unicode is the default nowadays everywhere. But then MySQL...)
Anyway, my CEO (also an American incidentally) was trying to keep me from it because he thought it's not high priority. However, we're about to go live in a French-speaking region, but which also has other indigenous languages (and therefore names), with their own "special" characters (I put "special" in quotes because for those languages, they're not "special" at all -- but I guess you get my gist by now).
Also, in previous jobs I have converted legacy systems to unicode and know what a pain it is down the road. Not to mention all the hard-to-find bugs if you don't do it, because some strings don't compare as they should, or people are just annoyed because their name is not shown correctly.
So I went ahead with the conversion anyway. We may never know for sure, but I'm convinced that I saved us some major customer frustrations, days of bug hunting and weeks of converting everything later, when existing data would need to be migrated.
So please everyone, just use UTF8 or some other unicode variant from the get-go. The few bits you might save otherwise are just not worth it.
I've been writing code to clean up a 2013 database dump. The database stored everything in LATIN-1 fields. Not because the data is in LATIN-1, but because LATIN-1 will accept any byte value. This makes error messages during input go away. See this bad advice on Stack Overflow.[1]
Some of the data is ASCII. Some is UTF-8. Some is Windows-1252. Some data is none of those, but is mostly ASCII except that there's a 0x9d once in a while. (Still haven't figured out what character set that is. From context, the ™ or ® symbol is intended.) So I have recognizers for these cases, and convert everything to UTF-8, testing every field value individually.
One column has garbaged non-English names. Someone had tried to "normalize" UTF-8 to lower case by using an ASCII lowercasing function on UTF-8 stored in a LATIN-1 field:
KACMAZLAR MEKANİK -> kacmazlar mekanä°k
Anita Calçados -> anita calã§ados
Felfria Resor för att Koh Lanta -> felfria resor fã¶r att koh lanta
I have the un-"normalized" form and can fix this.[1] https://stackoverflow.com/questions/44251813/unicodedecodeer...
There are a lot of Unicode-hostile environments out there. Java is old enough to always require explicit encoding declaration for pretty much any tool ... compiler, documentation generator, etc. Forget it at any one point and you get garbage. Reading or writing text files should always make the encoding explicit, but rarely does so. C#'s string methods all support, but don't require, a Culture parameter, without which you're practically guaranteed to do things like case conversion, or substring searches wrong in the general case. There was an awesome and long answer by tchrist on SO once about what the Perl boilerplate is to properly support Unicode for many or most circumstances (it's complicated and long and I doubt many people are going those lengths).
Point being, even when using something that supports Unicode well, the programmer still has to care, simply because text and language are messy things and it simply isn't possible to have a magic bullet that does everything right.
Much like the printing press, I'm 100% certain that the computing (and the internet specifically) is altering human written language across the world.
It is just so much easier to avoid anything outside ASCII because you can be certain ASCII will always work - even though some awful MS Access -> CSV -> SQL -> SQL -> Excel -> SQL -> COBOL ETL pipeline. No matter what version of any software is being used.
Technology has always shaped written language and we should fight to do better but at this point it seems inevitable.
(To be clear: I'm not saying this is a good or desirable state of affairs)
Eg A Ą Å Æ Ä are all different letters in most languages, not simply a pronunciation guide. I think most European languages use at least two from that list.
Manpages aren't written in TeX though (apparently they use something called "roff"), but they also contain things like `read' instead of 'read'... perhaps manpage writers tended to like TeX too?
http://www.read.seas.harvard.edu/~kohler/class/aosref/ritchi...
"Programmers use the grave accent symbol as a separate character (i.e., not combined with any letter) for a number of tasks. In this role, it is known as a backquote or backtick."
paul@tal:~$ od --format=x1z tmp/tonos-oxia
0000000 74 6f 6e 6f 73 3a 20 ce ae 0a 6f 78 69 61 3a 20 >tonos: ...oxia: <
0000020 20 e1 bd b5 0a > ....<
It's super fun when you are the only tech person in a Classics grad program and everyone else is turning in papers that look like ransom notes, because every fourth vowel has a different x-height from all the other letters. :-)Here are the two Unicode ranges (PDF):
modern Greek: http://unicode.org/charts/PDF/U0370.pdf
ancient Greek: http://unicode.org/charts/PDF/U1F00.pdf
Then again, try explaining to tech people the difference between ή and ἠ and how you can get ἤ or ᾔ. :-)
How do you enter them nowadays by the way? In the early days of the internet before unicode there were special fonts from SIL for example that first of all were using Latin characters (you'd type W and it would look like Ω) and secondly I think had the diacritics as separate characters, so the font would combine them. It was messy but at least you could type it on a normal keyboard. You could even read it in its ASCII form, most of the hacks were pretty reasonable.
Personally I type Greek using a vim keymap file, usually when writing LaTeX. I believe my keymap file is influenced by those older non-Unicode fonts, because I type w for ω and ;h for ή and >~h| for ᾖ. But my fellow students would mostly use Word. I don't know how they would type the letters, but commonly they would mix the tonos letters with everything else. I think what was happening is that Word would automatically substitute fonts that offered those codepoints, so it would end up showing Times for the letters with tonos and Palatino for the others (or something like that). Hence the ransom note effect.
Source: have a hyphenated last name. "No special characters in this field".
Ups and FedEx charge shippers ~$15 for each instance where they have to address correct.
So, yeah, the period thing is dumb, but automated correction is hard. Even experts, like SmartyStreets get it wrong often.
Japanese people can't have middle names (the citizen registry doesn't allow it), but foreigners can, so many systems will reject spaces in the name field. Meanwhile they're meticulous about making sure your name matches your ID exactly, leading to the situation I had.
The bank's web signup form disallowed spaces in your name, so I wrote my name FIRSTMIDDLE. Then when they processed my application they sent me an email "Your name doesn't match your ID card! Please approve the change to 'FIRST MIDDLE'."