Your last name contains invalid characters (2010)
blog.jgc.org
blog.jgc.org
One day, I got a call from IT, almost apologetically asking if they could change my email address because it was causing them issues. I pointed out I'd requested such approximately every 3 years since I started.
It turns out they'd never received those requests. I'll give you two guesses ..
It took a bit of time between Microsoft making the temp internet a 'special directory' and AV working properly in the directory, and the browser not saving attacker controlled literals to the filesystem to get past these. Also fun in NTFS because it could cause bluescreens was the AUX, CON, PRNT filenames.
Your post just reminded me of that kind of issue.
Surprisingly, the random text in the program was interpreted as valid program code. I was too young to understand exactly what had happened at the time, but now I understand it's because one of the valid forms of .com programs is a headerless chunk of x86 code for DOS, and I guess that website's output just happened to (a) not immediately crash and (b) invoke the DOS service for printing.
Both have better consumer frontends now (although I remember Deutsche Bahn still recommending this system around 2019), but those systems ultimately get all their data from this one, as far as I'm aware.
(Which also means it's extremely difficult to determine if a .com file really is an executable--there's no signature. It either decodes or it doesn't--and most bytes decode correctly because you want to pack the commands in as densely as possible. Things which will not decode are packing inefficiencies.)
(And back from the Z80 days I remember very carefully crafting assembly code that could be embedded in a BASIC program without causing it to puke. Some commands were unavailable and some values were not permitted--amongst them, zero.)
(Replace bar with passwd - I was blocked from submitting my original comment by Cloudflare's WAF...)
I thought me having a hyphen on my last name was bad. I'm glad no place I work at has tried to add them. This is a great ending to this story btw, and maybe a reason to try and contact them via phone if you never hear back at all.
For example, typing "@O'Shaughnessy Demo App" into Teams to fire off a command to our bot would automatically stop populating the autocompletes after the apostrophe is typed. So users would be forced to scroll through the entire list of usernames that begin with O to get to our app name to type a command into our bot.
Our workaround was to rename the application to "Demo App by O'Shaughnessy". Microsoft is aware of the bug after we posted in their dev forums, but has not fixed it yet.
You'd think that typing the full "@O'Shaughnessy Demo App do thing" into the app without relying on autocomplete would work, but it does not.
"Enter your name"
-Okay here's my name
"No, you need to enter your name in kanji. Of course, everyone has a kanji name, right?"That, along with refusing half-width roman characters, requiring them to be full-width and all that kinda shit. Basically stating up front on their own website that they suck at programming.
> But Japanese banks at least has a system which mostly solves at least those problems - everyone's name is written in Katakana.
That's workable, though I still remember seeing banks with comically short character limits for names.
I complained about this to my bank when I found out they stored passwords and their response was "don't worry about it - you aren't responsible for fraud".
I'd much rather live in a world where banks are idiots with passwords and I'm not responsible for fraud, than a world where I'm expected to be not an idiot with my passwords, and I were responsible for fraud.
Every day, I count my lucky stars.
How? By forwarding me his unredacted application chock full of PII. When I pointed out that most people probably wouldn't appreciate their SSN and more being shot around in emails by the management company, I was simply told it was "standard practice" for them. Following insinuations that their "standard practices" might just be "fucked" were summarily ignored.
Have fun trying to dispute an atm card transaction. The agent actively dissuaded me from doing it and said it was not possible.
Invalid: abc123 Valid: ab12cd34a4f9
Obviously, no uppercase or ANY special characters like “.” Allowed…
Everything is bloated with security theater, where the most critical things are vulnerable. For example, I saw a form for 2-Chōme requiring a phone number with three input slots. Entered my phone, it’s already been used. That’s fine, I just shifted one number from one slot to the other.
80-1234-5678 (already registered)
80-123-45678 (workaround to use the same phone number on a different account).
It worked.
Do they actually block hiragana and katakana? If they do that's probably grounds to sue.
>refusing half-width roman characters
With maybe the exception of arabic numerals, Japanese is nearly always written in full/monospace width. This is not unlike how English is nearly always written in half/proportional width.
When in Rome, do as the Romans do.
I don't have evidence at the ready, but I remember interacting with websites which complained about receiving katakana for both the name and カナ field for your name.
> With maybe the exception of arabic numerals, Japanese is nearly always written in full/monospace width. This is not unlike how English is nearly always written in half/proportional width
Yes, I'm fine with them using it. But they could put in a _little_ bit of work for me and convert my half-width characters to full-width, as the good websites here do.
If the field wanted only either hiragana or katakana (remember, it wants a simple reading guide) and complained, I'm not necessarily surprised.
What do you do at that point?
It's not my problem they have to deal with malformed data if that's all they will accept.
many websites in Japan has two name field: one is kanji, another is katakana.
https://shinkabukiza.pia.jp/membmng/RegisterNormalAction.do
Take this for example, the field in first row is full-width kanji, second one is full-width katakana.
The フリガナ (furigana) field is to indicate how the name given above is read, because kanji can have many readings including completely arbitrary ones. This also serves to indicate how computers should index and sort the names when storing and processing them, so it's still applicable even if the name is all hiragana or katakana and immediately obvious.
Furigana is also used to indicate how to read kanji in ordinary text, oftentimes when dealing with rare kanji or special readings, when the text must be comprehensible by everyone (eg: emergency bulletins), or when the text is written for people learning Japanese (eg: school textbooks).
Furigana is usually written using hiragana, so the reason 全角カタカナ (full width katakana) is specified instead of just full width is to inform the form's filer that he shouldn't write hiragana like he otherwise probably would.
https://en.wikipedia.org/wiki/J%C5%8Dy%C5%8D_kanji
Is it common to come across names using kanji outside that range?
Even in English this happens, look at all the variations of Robert (Rob, Bob, Robb, Robbie, Bobbie, ...) or similar names.
That doesn't quite grasp the magnitude. We're talking on the level of writing "Charlie" and reading it as "Alexander". Why? Because yes, that's why.
Providing readings using furigana is the mechanical answer to a very human problem.
I'm thinking.. I've only ever seen someone's name, will I make serious faux pas if I mispronounce it with the regular reading when I meet them?
(this is less pronounced in English because most of the time you can tell how to pronounce it by reading it. My last name suffers from the fact there is multiple possible pronunciations, so whenever I meet someone who has never heard my name they always stumble and look at me for help)
As for business cards, yes, if they have some non-obvious reading they will have kana or romaji there somewhere. (Usage of romaji on business cards is wider than you might expect, since it's seen as kind of the equivalent of a modern sans-serif logo for a business in some circles)
e.g.
Card with furigana: https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2F...
Card with kana elsewhere: https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2F...
Card with romaji: https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2F...
On the first example -- what is the significance of the character which is a circle within a circle? I'm assuming it's some sort of graphic or punctuation, rather than a kanji...?
The heading of that corner is just the pronounication of "m-take design" written out in katakana. The first line under the heading talks about the type of products the person works on (direct mail, leaflets, pamphlets). The second line talks about their specialities, as in the industries they focus on (cosmetics, health food), and then the last mentions they'll also do logo design, homepage design etc. I guess they wanted to emphasize the second line.
You'll sometimes see kanji-sized single circles used in Japan to indicate omitted characters (and you might see them on business card _templates_ as a sort of lorem ipsum, but unlikely on actual business cards), but this double circle doesn't have any specific meaning as far as I know.
The half-width/full-width distinction is a hold-over from the DOS era when all fonts where monospaced. It doesn't have a place nowadays and should have never be included in Unicode, just change the font or use a better text rendering algorithm.
Nah, it's probably some Japan-only ancient mainframe OS with Shit-JIS over EBCDIC.
"Hmmm, your ID card says 'Christian Joseph Anderson' but our records show 'Anderson Joseph Chri', are you trying to pull a fast one on us???"
In my experience computers are the bigger issue. A human can think outside the box in a way a `strcmp` can't.
Computer says no
https://www.youtube.com/watch?v=0n_Ty_72Qds
---
A human can do many things, but never underestimate the power of the system (Hello Moloch) to remove said creative abilities from humans. Also never underestimate how creatively stupid humans can be.
Middle names are less common where I live. The locals tend to add my middle name to my fist name and then complain that it's long and doesn't match their database at first glance.
interesting anecdote
that being said, i think most of these websites that ask for a kanji name will require you to show ID with that name when you show up in person, so you might run into trouble if you just pick random characters.
Conversely, I've had to fill forms in Korea that included space for only 4 letters, including both first and last name. Needless to say, my 3 first names plus last name didn't fit.
This caused big problems whenever he booked airlines tickets as the name on the ticket didn't match his ID.
The initial happened to match his father's given name, and when he was old enough, he reliably chose to take that name as his middle name. It is certainly an endearing story of filial devotion, and a distinct lack of finicky SQL databases or web input validation in the 1950s.
I wonder if we know the same guy. There can't be many Js running around.
Moved to UK last year. Couldn’t put my English name in anywhere and had to use my full name. Feels super weird getting called so.
Could be apolitical protection against chess grandmasters: https://en.m.wikipedia.org/wiki/Wesley_So ;)
I don't know if it truncated it automatically or it just stopped accepting input after 20 characters and I of course did not notice since the password entry fields were masked.
My last name end on "ić" so every time, even on many EU shops and websites, I get to choose if I want to write my name wrong with "ic" and definitley get the package(cause the post understands) or write my name properly and get my package in 90% of cases.
Its just plain weird that a Spanish/French/UK online shop will support their weird characters but not full Unicode.
every time i walk into a coffee shop i have to judge how much trouble spelling my name 'correctly' will be and will switch to "Y" in those cases
like, I thought Joachim was supposed to blaze the trail for the rest of us soft J's :(
Between plain 7-bit ASCII and full Unicode there were a few decades of make-do with several 8 bit encodings.
A couple of years ago when I wrote a little webapp to do some scheduling stuff for our community's yearly conference, I just used plain unicode strings for usernames, including spaces. I didn't really see why not: The database (SQLite) handles unicode, the backend (golang) handles unicode, the frontend handles unicode, nothing ever gets passed into a shell program or put into CSV or anything, all queries are parameterized; why bother making pointless restrictions?
My last name contains a naughty substring. I feel for the author and their hyphen!
Which, come to think of it, is kind of a funny notion, and a nice reminder that despite our differences we really are all related!
I can vaguely imagine how someone would _think_ its a good idea (it's not, to be clear) on a website where users might see other users' names, something like CRM SaaS. I can't understand how anyone would think that validating users' real names on a purely customer-facing website is a good idea. Maybe frustrated customers (can't imagine why) have names like FuckVerizon on their account?
Well, I don't live in Sweden ... so right now every system is presenting me as Aron Gunnar Mauritz, making every person interacting with those system refering to me as Aron Gunnar Mauritz. Which is madness to me.
Naming customs are subtle and hard for other cultures to grasp.
I'm confused, Isn't EU much more privacy focused and with GDPR etc it's easier to reduce your internet footprint?
[1] https://en.wikipedia.org/wiki/Principle_of_public_access_to_...
[2] https://en.wikipedia.org/wiki/Swedish_Freedom_of_the_Press_A...
EDIT: There are even search engines to find out who in an area have been convicted of various crimes.
I'm from a country that _has_ middle names but doesn't really use them, so I have to choose between "is this likely to be compared against official records that will need the middle name, or should I omit it so the system doesn't start referring to me as "Hi First Middle".
[1] https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
My case was actually this.
I took a call from a person with a name that was one word. Not first name, not last name, just one word.
I never found them in the system as the name wasn’t in the database as far as I could tell. They were quite irritated.
People can have multiple full names.
Edit: Nope, rule 40.
Famous examples but for stage names are Madonna or Cher.
I knew someone who had one that has a mononym and would regularly run into issues.
I believe they did get their passport issued with their mononym, but it caused all sorts of problems with international travel. Some airlines would require duplicating their name for first and last, then cause issues during check in, security or immigration, because of the mismatch between passport and ticket/visa/etc.
I do.
The idea of a plane ticket for ‘Cher Cher’ is quite funny, though presumably wasn’t for her.
I suspect you are applying just a variation of the same every-name-must-follow-my-arbitrary-rules game. You're trying to correct the assumption that everybody has of a first name and a surname with the equally wrong assumption that everybody understands an English word.
One counter-example for "People have, at this point in time, one full name which they go by." I learned from the 2020 election concerned the Georgia voter fraud claim that someone used the name "James Blalock" to vote, years after he died in January 2006.
Thing is, it was his 96-year-old widow, "Mrs. James Blalock Jr.", who voted. That was the name she registered as, back when using "Mrs. <Husband>" was a common custom.
Once, checking in to a flight within China security didn‘t want to let me through.
Because „Müller“ (as an example) on the passport is not identical to „Mueller“ on the ticket.
I eventually made it through when I pointed out that the machine-readable code on the bottom of the passport contains the string „MUELLER“.
Names outside US-ASCII are always problematic.
I googled "Chinese passport" to have a look, and conveniently for them, they have the ASCII version of the owner's name under the Chinese name.
I did pick another name at that time because I had been waiting for an hour, so practical concerns trumped principles, but it's ten years later and I still feel offended when I think about this.
It has literally never been an issue, and I suspect the agents are just used to this.
Of course, but, to give you a sense why I felt so annoyed, the conversation went something like this:
> What is your last name sir?
"de Bakker"
> That name is invalid. What is your name?
My last name is "de Bakker".
> That name is invalid. You need to give a valid name.
"de Bakker" is my last name. It's on my passport.
> Sir, if you don't cooperate I'm gonna hang up and proceed with the next customer. What's your name?
So it was very clear that I was the one at fault here.
European websites never bother me about it, of course :)
Doesn't bother me that much but I've had moments where a service told me that "Your name MUST match the name on your ID", only to then complain that I can't use a certain character. Which is present on my ID. Well.
The postman understands "Kbenhavn", "K~#benhavn >&" and so on.
Very occasionally some address verification thing can't match my address, as one side or the other within their system is corrupting it.
In a just world, the website would be taken offline at once, and kept off until fixed.
Plenty of people do have spaces in their surnames (like Conan Doyle) and shouldn't the system be able to distinguish between "Conan Doyle" and "Conan-Doyle"?)
I eventually moved, and every state after that wasn't quite that dumb.
Needless to say, this matches nothing at all, like airline tickets, or my birth certificate or passport. Shockingly, this was never a problem.
Why is this mandated by the government?
"VLKK Chairperson Violeta Meiliūnaitė claims that legalising such a spelling would “violate the Lithuanian name system and have a negative impact on the country’s linguistic and cultural identity and distinctiveness”."
https://media.efhr.eu/2023/09/27/vlkk-legalisation-of-female...
https://en.m.wikipedia.org/wiki/Naming_law
Hopefully someone will add the reason for Lithuania.
Iceland:
Parents are limited to choosing children's names from the Personal Names Register, which as of 2013 approved 1800 names for each gender. Since 2019 given names are no longer restricted by gender. The Icelandic Naming Committee maintains the list and hears requests for exceptions.
New Zealand: Below is a list of banned names in New Zealand:* [Asterisk], 4Real, 89, Anal, Bishop, Constable, H-Q, II, III, Justice, Justus, Knight, Lucifer, Mafia No Fear, Minister, Mr, Queen Victoria, Royale, Saint, Sex Fruit, Talula Does The Hula From Hawaii. Note I suspect * is special because I think it is a is placeholder in the Dept. Internal Affairs for none (I had an acquaintance that changed their name to a single word and the asterisk appeared in their surname field officially - also see http://wookware.org/name.html )And hilariously Australia had this girl called Methamphetamine Rules recently: https://www.theguardian.com/australia-news/2023/sep/19/can-y...
Bus shelter 59 (or similar) Hitler Benson & Hedges
Polish does this for some local last names, mostly the ones ending with "ski" (they end with "ska" for a female)[1]. This makes grammatical sense, Polish adjectives change their form depending on the gender of the noun they apply to, and those names are kind of sort of adjective like.
Czech goes even further and applies grammatical rules to all names, even foreign ones. Czech news broadcasters will literally say "Melania Trumpova" or "Michelle Obamova".
Incidentally, those gendered forms are a pain in the ass to deal with in user interfaces, particularly if you don't have gender information for your users and/or want to support nonbinary, something slavic languages are really not designed to do.
[1] The US doesn't enforce this rule of course, and therefore it's not unusual to meet a female American with Polish roots with the surname "Kaminski" (or sometimes even "Cumminskey").
Also it's a common choice for translators of books written in those languages to give female characters the masculine versions of their surnames so readers won't be confused about who is related to who. Perhaps an extreme example is that some English translations of a particular Tolstoy novel have the title Anna Karenin.
It was around 2010 and I joined (as a contract) company which used Google's Cloud Platform for their products. It was initial phase, but stuff was ramping up quite quickly.
So, since I was contractor, and it was supposed to be short gig, access was given to my Google account, which was one of not that many places that had my full name visible (mostly for visibility and official matters).
Week or two later I was to do something with CDN, or Cloud Files, don't remember the product, but it was managed through Python script. As soon as I logged in and tried to do something I got familiar error about unsupported encoding.
That was weird, but as that was Python script I quickly figured out that the reason was that Google was pulling my name from their account and their script couldn't handle it. Oh well, happens, I mailed support (what a time to be alive back then!) and went toward my way not really expecting much.
I received response shortly after, that in this case it wasn't possible to fix the access script. After few back and forth where I pushed against using latin chars in every Google product possibly support finished with something along the lines "well, you can just change your name in that case".
Funnily enough, for company that was hiring me it was the last straw and they took their business elsewhere. I, personally, was annoyed for a long time, but today it only makes for a good anecdote. I never used any GCP products though.
Last name contains invalid characters [2010] - https://news.ycombinator.com/item?id=27769788 - July 2021 (2 comments)
Your last name contains invalid characters - https://news.ycombinator.com/item?id=1438355 - June 2010 (73 comments)
Mistakes and technical limitations are understandable, but why blame the customer? An automated error message was still written by a human. Human people decided the response to "invalid characters".
Subbing in a computer for a human doesn't create a valid excuse. Saying, "being nice doesn't scale", doesn't make things right.
My war story- tried to get a copy of my birth certificate. My mother doesn't have a middle name. Never did. First and last only. The county office required me to tell them her middle name. After much arguing they said her last name was her married name and her middle name was her maiden name. When I asked what her name was before marriage, I finally got it much to the dismay of the employees.
At least I had a human to look at. That eventually, with much work, used common sense over programmed logic.
There should only be one input field that asks for a name. Even systems that require a legal name should be like that as there are people whose legal name cannot be written in two forms.
(p.s. I did present what I wrote above on said mailing list back then, but as far as I remember people simply decided to ignore it as too rare to care about, in addition to those who for some reason taught I just made it up (why should I?))
This is the hyphen-minus, U+002D: - (and ASCII too)
This is a hyphen, U+2010: ‐
This is a minus, U+2212: −
IDK but every time I programmatically create an AWS resource, I have to go crawling through masses of disorganized documentation to see what character substitutions to make.
Hyphens do destroy AWS's database. :shrug:
Have one input: "What should we call you?"
and, if you really need it: "your full name"
> Ask HN: How to handle Asian-style “Family name first” when designing interfaces?
the discuss shows it's so hard to handle name of people from any countries.
There are cases where that happens for real and not just due to other computers not supporting Unicode properly; for example it is completely unreasonable for every postal system in the world to have to support every language in the world for addresses.
But even beyond that, even expecting every postal system in the world to have to understand every script is itself not feasible. I don't mean this to offend, but Arabic is just a scribble to me. I know from reading articles about the difficulty of typography that Arabic letters are generally strongly affected by their preceding and/or following characters, but that only from HN, not from my normal day-to-day experience. I couldn't even parse it into letters correctly or safely. The ideographic languages have no spaces, so I can't break them into words safely. (After some non-trivial study of Japanese, I could mostly do it, but still only mostly.) It's not a reasonable expectation to put on every postal system in the world.
I.e. I don't understand why these problems mean that a postal system might elect to not support all scripts in storage (I get why it means the OCR software for a postal office might not support scripts outside their own country's scripts).
This is much like the naming scenario, where I may need both your legal name for legal reasons, which may have local restrictions on it, and I may also have a field that is basically "What do you want me to call you?" which may have arbitrary unbounded Unicode for all I care, I don't care if my system calls you the Lord and Master of Zalgo Text. But there's no programmatic way to get from the latter to the former, or indeed, even a human way.
Surprising number of systems can’t handle this gracefully even when they handle McGuffin as a last name just fine.
Order a ticket online, they pick the transliteration for ø as o, passports transliterate ø to oe - the less ambiguous choice. This generally isn't a problem until you want to travel to one of the APIS countries - USA or UK, which won't let you check in if the name on your passport doesn't match the name on the ticket.
I have tormented many check-in counter staff with this, especially at regional airports that don't see a lot of international travel.
it should be possible to file a complaint to you country's standardization agency and they will pressure the airline with fines
This is always a paranoid worry of mine. Even worse, some airlines tend to include the academic title in there too, meaning I've had tickets reading "MRDRWRONG, PROBABLY".
Not only does it not match my passport, but it also makes it look as if I'm a murderer.
We had to do it because the underlying backend is Spanish and the APIs are insane. Behind our JSON-wrapper is an adapter to translates it into teletype friendly format … spaces, new lines and tabs makes it look like a ticket you’d get from a travel agency in the 1970s.
This format is also the reason why your name is truncated on your ticket (or so I was told). IT in the airline industry is insane and many systems are many decades old at this point.
A bit surprised you have trouble with APIS though.
This is why I don't use Namecheap anymore fwiw. They used to accept 'ø', they redesigned their interface, I had to add a new debit card to my account, their new system didn't accept 'ø' yet told me to write the name "exactly as it is on the card". Customer service just told me to write a name that's not mine in the billing info; I moved to a different registrar instead.
Why do we even have that field anymore?
Also, to the original article's point, it looks like these websites would exclude like half of Québec's population; dual last names are very common here (women keep their names after marriage, making the "mother's maiden name" security question pretty stupid).
Diactrics on uppercase letters weren't that common before: they were officially supposed to be there on uppercase characters but:
- old french typewriters didn't have them (they had the lowercase é, è, à etc. though)
- people using early computers had no idea how to encode them at all (but é, è, à etc. were on the keyboard) [1]
But "ATTENTION MARCHE" and "ATTENTION MARCHÉ" are two very different things. It's simply wrong to use the first one if you meant the second and vice versa.
Now I know a friend who lives in the US and who simply dropped the diacritic from its family name: so now he and his kids have a "different" family name than his family in Europe.
He found it easier that way.
[1] People using Word in the early nineties had to type 'é' and then select that character and go to a menu to transform it to uppercase, because "shift+é" would enter the digit '2'.