It's not as big an issue as it used to be, at least. Before I've had online transactions failing because of a mismatch between my name (with ü), and the name on the card (with u). The systems seem more forgiving now, having handled that case or something. I also remember being a bit scared traveling to Japan many years ago, as we were told it was SOoo important that the names and everything matched to gain entry. And then the name on my ticket was completely mangled. But no one cared.
Here's a SO post about someone with the last name Null: https://stackoverflow.com/q/4456438/923847
I would have thought that people from CJK-countries were more understanding of encoding-to-latin weirdness than most, but apparently not.
It's much easier to identify mojibake (they tend to be extremely obvious in CJK encodings) than to remember canonical spellings and other variations in a whole bunch of different languages. Airport staff probably know that "oe" and "œ" are interchangeable, but that's about it.
German is more complicated though with all the substitution rules.
Not to mention Germans who actually have an ue in their name, still pronounced as ü, but written as ue only, never as ü. Or someone may be called Gross, but it would be incorrect to write it as Groß, while someone else's name may be Groß with the acceptable alternative spelling Gross when ß is unavailable.
Not in all cases. In Germany and Finland (maybe all EU passports???) ä is spelled ae, ö is spelled oe in the machine readable part (umlauts shown in the "human-readable" part). This is important to know when you need a visa.
For Germans this is not a big problem because it has been like this forever if the umlaut is not available for technical reasons. For Finns this is a problem, because this "transcription" is completely unknown in Finnish. For a couple of weeks now it has been possible to get an electronic visa for Russia on the internet. Reportedly many Finns with an ä in their name (that's not uncommon) dropped the dots when applying for their visa, because an ä is not accepted. At the border they were not allowed to enter, because the machine-readable part of the passport has ae instead.
I do wonder what happens to ű and ő though.
https://www.icao.int/publications/Documents/9303_p3_cons_en....
Ü is written as UE, UXX or U
Ű is written as U
According to https://en.wikipedia.org/wiki/Machine-readable_passport#Name... Hungary uses UE for Ü, but there is no reference given. According to the same article Russia uses even 2 different transliteration systems depending on the type of document.
My experience has actually improved substantially in the last 10 years or so, and most of the government systems I encounter these days actually handle it properly (as well as handling suffix properly too). That said, I've somewhat recently started having trouble checking in for flights again -- I flew last month and it took the ticketing agent >20 minutes to find my reservation on both the outbound and return flights, even despite my providing the 'confirmation code' / itinerary email (we were checking bags & flying with infants, else I'd have done online check-in).
It can be really frustrating -- though I'm hopeful it will continue improving and hopefully be a smoother experience by the time my kids are adults.
That being said last time I went to the US the person booking the ticket swapped my first name and last name. Only the person at the baggage dropoff noticed it, and after much deliberation they suggested to leave it that way. I went through with no issues apart from not being able to register the mileage.
Similarly, automated check-in kiosks are then usually unable to find the reservation via credit-card or passport scan -- meaning you're back to looking up the reservation code, and even that often fails, as if the apostrophe just flat-out causes issues with the query/lookup or something.
It can be very frustrating, and I'm increasingly often impressed (and vocalize the same) when I spell my name and the agent enters it correctly AND the system flawlessly handles it, too! The DMV systems in my state are one such example where I used to have issues but, in recent years, the problem appears to have been wholly addressed/handled.
Ugh, yes. And it's insane how many people seem to just NOT KNOW what an apostrophe is.
> checking in for flights
Yea, airlines seem to be one of the worst offenders. I have Precheck but Spirit in particular is never able to match the name on my ticket to the name in the gov't database so I never get it. Just one more reason to avoid flying them I guess.
What is worse is people who "fix" my name by moving the second half of my first name and making it part of my last name. I'm an adult. I know what my name is.
However, now that I've moved to the United States, it's been a bit of an annoyance.
The letter e with an acute accent causes all sorts of UTF-8 encoding issues with many services, not just airliners. If you interpret the UTF-8 é (0xC3A9) as ASCII it becomes à (0xC3) + © (0xA9), so my name often comes out as 'Léon'.
Airlines make it worse, because they strip both characters during sanity checking, so my name comes out as 'Lon', which has caused me problems a couple of times as the name on my passport did not match the name on the ticket.
As latin1 (ISO-8859-1) or Win-1252; ASCII doesn't have either à or ©.
latin1 is the default for text, including HTML, if you don't specify in protocols such as HTTP (modulo some stupidity from the WHATWG where it might be Win-1252 instead) and Windows-1252 is the default encoding in Windows in the USA (at least, prior to the Unicode APIs being added. The old APIs probably still exist though…). So these codecs pop up a lot in places where people who don't know what they're doing end up touching text.
If the transport, content-type, lack of charset declaration, and sniffing fail to determine an encoding, both specs use defaults based on the configured locale, for English that's windows-1252 [WHATWG: 12.2.3.2 W2C: 8.2.2.2]. latin1/ISO-8859-1 is prohibited. [WHATWG: 12.2.3.3 W3C: 8.2.2.3].
It is amazing to me where I see failed encoding like that. For instance, many SEC filings and job ads for tech companies. I mean, I feel like I'm expected to spell things correctly on my resume and emails at work...
e◌́ => é
Which should keep the `e` intact, while the combining acute accent (0xCC 0x81) may "only" get converted to a `Ì` which may be stripped. 0x81 is undefined in Windows-1252, so I have no idea what would happen to that, but probably be stripped as well, keeping just Leon.
http://i.imgur.com/4J7Il0m.jpg
What these things all reinforce is that a lot of programmers take text encoding as a given, and don’t realize all the potential places for errors to sneak in.
Also seeing that with accentuated uppercase letters in French, even in nouns, because it's hard to type them on Windows.
People still use accents in lowercase of course, but think that it's incorrect to use accents for uppercase letters, even when handwriting.