Falsehoods Programmers Believe About Names
kalzumeus.com
kalzumeus.com
I go by my middle name (my parents thoughtfully gave me my father's first name and even middle initial, which has caused no end of confusion for many a credit bureau over the years). Unfortunately, the State of Indiana's birth certificate system assumes beyond any possibility of override that (1) everybody with children has a first name, (2) that first name has no spaces in it, and (3) the middle name is insignificant. So my kids' birth certificates have my dad's names on them as the father.
But hey, who am I, a mere parent, to say what my name is?
The real world is full of organic detail that is difficult or perhaps impossible to capture in full in a software system. A name should be just a string. If there are business requirements for sorting or name-of-address (e.g. "Dear Mr. Jones") they should be done with heuristics, with human intervention invoked in sticky situations if necessary. Sure, it's difficult, but you know, a hundred years ago that stuff was all done on an individual case-by-case basis by human beings; if your business assumption is that it can't be done cheaply enough without eliminating human intervention, perhaps you should rethink your assumptions.
Otherwise, let me assure you you're irritating every customer whose name doesn't meet your arbitrary rules - in exactly the same way that Google pisses off everybody who needs human support. Everybody thinks that sucks. So at least you should get people's names right, using human intervention if you have to. (And addresses, too, but we just had an extended thread about that, like last month.)
If the distinction matters, then whatever system is in place had better be prepared to use my ordinal (III) as significant, but had better not put it as part of my last name.
So I will go by: "A-"; "A- Z-"; "A- H- Z-"; or "A- H- Z- III". I never go by "H-"; "H- Z-"; "H- Z- III"; or "A- Z- III".
So, for a while, I lived on Newp0rt Way.
Indeed. The real question is, why are you asking for their name in the first place?
Consider these three scenarios:
(a) You are an airline, and the name you collect has to match their government-issue ID
(b) You need to mail something to the user
(c) You wish to create an account for them on your blogging website
The requirements for you to understand what they are trying to tell you their name is very much depend on the situation (context matters, as I need to have tattooed on my forehead). For (c), screw what is their first name, or last name, or middle name, or whether they have such a thing. For (b) the only requirement is "write something down what the post office in their country will understand" - again, never mind whether you understand it. For (a), how is this going to be matched by TSA? Do what is right to conform to the API.
Even in the US, people have very complicated names - here is a not unusual name from my local paper's birth announcements: Kalamanamananaueikalani Tomoko Namakauluhonuamea'ililamalama Wengler-Ioane
Does anybody really, really need to deal with this in its full glory? It should come down to "What would you like us to call you" and "are you M/F" (if relevant).
1) A simple, ascii A-Za-z0-9_ identifier unique to program/service is not required to use my program/service.
1a) My program/service cares deeply about your real world name and all it's nuances and variations. It requires full and complete understanding of your name rather than just a simple authentication mechanism.
2) My program/service will correctly handle the worlds naming conventions past, present, and future to the exclusion of other features / actually shipping some day.
3) Instead of following 80/20 "rule", my program/service will names 100% perfect!
4) My program/service will cater to all languages/cultures/subcultures/niches past, present, and future rather than any sort of target audience.
Just read the complaints here about airliners screwing up people's names by assuming that everything is ASCII/LATIN1/MISC_CHARSET. That makes a difference when you get to the airport and they won't let you on the plane because your ticket and your id don't match.
It's also an easily solvable problem. Just include a 'name' field. Make sure that it's sanitized correctly and supports UTF-8 (make sure that your database is properly set up to use UTF-8 as well, or else all manner of problems can happen, least of all queries will get bogged down as the database does charset conversions for string comparisons). If you care so little about the user's name, then why bother to have separate firstname/lastname fields?
{edit} With respect to the 'other complaints,' I was referring to the comments on this story: http://news.ycombinator.com/item?id=1438355 (The story which is what Patrick's post was in response to).
Having worked in the airline industry, I can tell you, this is not a case of airlines making a poor assumption. Airline systems are generally incapable of dealing with ASCII, let alone Unicode text. They typically use EBCDIC. That's why you never see lower case characters in your name on a boarding pass.
That makes a difference when you get to the airport and they won't let you on the plane because your ticket and your id don't match.
Does this actually happen? Because given the fact that airlines have been incapable of printing names that don't can't be described by EBCDIC since, well, forever, I'd assume that they're willing to ignore cases where the name on your ID contains characters that their computer system literally cannot represent.
Do the screeners know to this? I'm willing to bet that a screener would deny you based on a name mis-match. What do the screeners know of EBCDIC?
As for TSA staff, it is important to realize that this EBCDIC issue is not a problem that only affects one airline: it affects almost every single airline, including all flag carriers that I know of. The airlines have a pretty tight working relationship with TSA; note that they send electronic messages to the TSA describing their customers when people book tickets and just before departure so as to give the TSA the opportunity to prevent them from flying. Given that tight relationship, I'm pretty sure that the TSA knows a great deal about the limitations of EBCDIC and has informed their staff about the limits of what a boarding pass name field can possibly represent.
Keep in mind that screeners will let you fly even if you have no ID on you.
Anecdotally, I've never had issues with tickets that had my partial name on them in the past either ("Matt" vs. "Matthew") although this may have changed now that TSA requires full names and birthdays (I always use my full name now).
A travel agent I asked about it once told that if the names match 80% or more of the character they assume a typo was made.
There are 2 ways to assign SKUs, sequentially (start with "1" and increment 1 for every new SKU) or not sequentially. I have never seen anyone do it sequentially. (Although some excellent systems keep 2 SKUs, one of which is sequential and is used as the primary key and for all indexing. This is the best way I've ever seen to handle SKUs that change, but that's another story.)
Almost everyone wants a smart or semi-smart SKU. So that by simply looking at the SKU, anyone can tell what it is without reading the description. You know, the first digit is Commodity Code, 1 for shoes, 2 for pants, etc. Then another digit for color, another for size, etc. This works well until you have ten colors; then you need 2 digits or alphas.
But wait, there's more. Let's put hyphens (or some other delimitter) between the product descriptors and the vendor data, manufacturer data, and customer data.
So now you've covered any possible product with your super slick smart SKU naming system.
Until something comes along that isn't covered. (Now we have military items with 14 other considerations.) So we come up we a second totally different scheme. Then a third. Then a 4th, etc. So now you can tell anything about an item if you know which scheme it falls under.
But wait, there's more. You should be able to enter any SKU into a form field regardless of Smart SKU scheme. (If the first digit is "9", then use Smart Scheme 3. If there's a hyphen in position 3, then use Smart Scheme 8, etc.) Your form logic should be able to intelligently guide the user based on the rules of the template or scheme.
I have built apps where users can design and build their own Smart SKU templates, which are then used to enforce compliance and guide operators. These have generally worked pretty well.
Is there some way to do the same thing for human names? I dunno, but now you got me to thinking about it. A combination of standard templates and custom templates oughta cover most possibilities. Some basic logic with optional pop-up forms which uses the templates as parameters should work. Something to think about...
32 opcodes? That's plenty! We'll never run out. :)
Not that URI:s are that great for showing to humans, but if we added some visualizations template plug-in system for parts of it (similar to how many web servers map URLs to renderings of data resources).
Or is this kinda what you mean with templates?
http://news.bbc.co.uk/1/hi/sci/tech/8206280.stm
Certain characters like the defunct yogh have had issues in Unicode -
http://en.wikipedia.org/wiki/Yogh
http://news.bbc.co.uk/1/hi/magazine/4595228.stm
Some names can be changed if they're cruel or too unconventional -
http://news.bbc.co.uk/1/hi/7522952.stm
"Professor Robert Smith? (the question mark is part of his surname and not a typographical mistake)"
"Is that you Professor Robert Smith??"
And I wonder if he introduces himself as "Robert Smith Question Mark" ;-)
(...although there are plenty of other names with exclamation points, due to languages like Xhosa.)
Particularly confusing to people when his passport has the name spelled one way and signed another.
Nobody has been able to get a straight story from him about why.
Then when they finish laughing, hack up some system to split the single name field that doesn't work very well.
(It gets even worse when you have unusual sorting rules that really only apply to Western names, like treating Mc and Mac the same.)
My customers said "Fine, alright, two sort features. One for Japanese, one for foreigners."
They were less than happy when I told them that foreigners haven't agreed on lexicographic sort, either. (To use one example I'm familiar with, in Spanish, "ch" is one letter, so Chisako comes after Consuela.)
Oh, there is a separate right way to order prefectures. (Which I learned about in an email from a coworker saying "Patrick, come on, use some common sense next time. Have you ever seen prefectures listed lexicographically before?!")
You've correctly identified a problem, but misassigned the responsibility. The problem with the single-name split is that it is impossible, full stop. When you forcibly try to do the impossible, you always get bad results.
You appear to be proposing that we force the users to enter first and last names, then we can trivially sort them by last name. But for the exact same reasons you can't write a name splitter after the fact, you can't have the user break their names into "first" and "last" either. You haven't solved the problem of there being "no last name" by changing your input form. You've just moved it around. Now it irritates the customer instead of you, which is often a bad trade, and sort by last name still doesn't work because customers for who a first/last split doesn't work have fed you one or another variety of garbage data. Garbage data which is now even harder to find; at least when you had a one-word name you had a clue that there was no first/last split.
Which of course goes against the spirit of the original article.
The receptionist SHOULD be shown a list but ALSO wants to be able to search for me if I don't show up in the list (and may ask for my ID/get a particular spelling).
(It was a comp'd room in Vegas, so I have no idea where the mistake was made, but this can't be a unique occurrence.)
There are four fields in the database:
first_name
last_name
display_name
primary_email
All 3 name fields are optional. The email address field is required, but you could use a customer ID or username if that is more appropriate for your app.The name are never accessed directly, except on the one form where you can edit these fields. In general, they are accessed by these two helper methods:
@property
def name(self):
if self.display_name:
return self.display_name
if self.first_name and self.last_name:
return self.first_name + ' ' + self.last_name
return self.primary_email
@property
def last_first(self):
if self.first_name and self.last_name:
return self.last_name + ', ' + self.first_name
if self.display_name:
return self.display_name
return self.primary_email
Then we do a culture specific case insensitive sort. This should cover most cases pretty well....... I hope :-)Obviously, you don't want to ask too much up front (to fight what I'm coining as "form-fatigue"; you know, where there's so many questions on the form that you just give up signing up), so you may need some decent heuristics to give you reasonable default values for these things that the user can override if they want.
As far as the Artist Formerly Known As Prince symbol lying on its' side with a happy face in the middle, in the credits: "I'm the storyboard artist formerly known as J. Todd Anderson. That's all I can say about that." It's a private joke between J. Todd and the Coens. Prince and the Coen brothers are both from Minneapolis. (Dayton Daily News: 3/22/96)
I can live with the romanized version. It's been in use for more than a thousand years and it's a bit late to complain.
http://en.wikipedia.org/wiki/Identity_document
It even has a nice check digit.
But even that breaks in some cases :) so Patio's point applies.
Still, it makes building local-only apps easier.
:)
(Hello, fellow Uruguayan)
And I am not even going to say sorry.
> And I am not even going to say sorry.
With comments like this, one comes off as a douche, which isn't exactly going to engender trust with potential customers/clients. Whether it's 'your' fault or not, how you deal with your potential customers/clients is what really matters.++ People's first name is their "family name".
++ People have first names and last names.
I may be mistaken but I believe this is true in Vietnam (and probably other countries as well).
In the western world, we assume that someone's family name is their last name. In countries such as Korea, people assume that someone's family name is their first name.
Either assumption can be wrong if you're dealing with international customers. Some people might not even have a family name.
The most common ones are probably the easiest to deal with...long names, only one name, punctuation in names. Just relax your validation requirements. We just took care of 99% of names not common to western culture. But I don't think if you are designing an application for, say, the Olathe, KS youth rec sports registration website that you need to be too concerned with folks who have names that have characters not mapped in Unicode. If a person's name is so unusual that it doesn't fit even culturally relaxed input requirements, then they've probably already dealt with that problem before.
Many Japanese people think racial/cultural homogeneity is practically a national trademark, and this issue has bitten nearly every Big Freaking Enterprise system I've ever seen in Japan. Do you think your global-facing web app or small American town is going to be less culturally diverse than Japan is? That strikes me as highly improbable.
And, like I mentioned, in cases for applications targeted towards a homogeneous group, for those that have names so far outside the cultural norms that even liberal standards cannot accommodate them, then they've probably encountered the same thing already before, and they've come up with a way to adapt.
Allowing a liberal range is like me going to France and asking for ketchup with all of my food. But having a really unusual name, like one with unmappable characters, is like me going to Riyadh and expecting them to serve booze with my meal. Should I be allowed to drink there just because it's acceptable in my culture? No. It's illegal there, and I just have to adapt whether I like it or not.
Again, I emphasize, what I am advocating is for applications with a culturally homogeneous audience, not global, multicultural apps.
Some of them, including at least one prominent politician and several people I represented in a professional capacity, are adamant that reversing the order is incorrect.
(Unsurprisingly, many Americans are unaware of this convention. Sadly, many of them persist in being unaware of it even when it is written on the meeting briefing and explained verbally right before the meeting starts. sigh Twelve corrections in 3 days -- my client was not happy for that trip.)
Does your hotel, car rental service, university, etc, handle this case correctly? We blazed a path of frustration through Chicago and Michigan last time.
Further fun: a Taro Yamashita born and raised in the United States (or any other Taro, for that matter), and present at the same university for the same conference, might request the exact opposite treatment! And if you spell his name as YAMASHITA Taro on the meeting agenda in defiance of his preference, you're doing it wrong!
I mean, wouldn't it be pretty chauvinist for me to get offended at it or demand that they follow Western name ordering? It feels like it'd be in the same category as complaining that they don't serve my favorite American food at the conference, or that the cars are driving on the wrong side of the road.
You are.
The canonical source for the correct way to spell/write someone's name is that person. You should never tell someone "No, you're writing your name wrong".
Between countries, I think generally the right approach is to use the country's naming customs, and adapt foreign names to them, to the extent reasonably possible. In Greece, for example, it's customary to transliterate people's names into the Greek alphabet, especially if the source is going to be read by Greeks (newspapers, etc.). The person known as "George Bush" in the United States is more commonly referred to in Greece as "Τζωρτζ Μπους", for example--- despite it being a rather ugly transliteration in this case, due to a bunch of the consonants not existing in Greek.
Are you really arguing that this Dr. Yamashita (or Mr. Bush) can tell them they can't use their own alphabet in their own country, because that's not how he likes his name written?
My middle name is a musk ox.
How far does politeness dictate that you have to go to accommodate me?
"Your name does not fit into a Western firstname/lastname
format with only ASCII characters, please choose another
name."
This is on par with expecting everyone to speak English just because you speak English. By the very setup of the form, you are implying that your way is the 'one true way' or at least that you don't care about people that don't fit into your pre-defined set of expectations (i.e. "Don't have a firstname/lastname? I don't care about your business! Your money is no good here!").How hard is it to just have a 'Name' field that supports UTF8? Sure it's not a 100% solution, but it's a lot better than the 20% (or less) solutions that we have out there right now.
So, yes, there are a few you can safely ignore. The problem is the set you can safely ignore is probably different for almost every job you'll have.
If your operating assumptions are true enough for a given purpose, there's no reason to change them. If they deviate enough from to generate problems, they will need to be replaced ... with other approximations (not with "the truth").
-- This is why apparently simple and "finished" programs alway need "maintenance" - even the most seemingly simple assumption need tweaking as the world's conditions change.
Nicholas If-Jesus-Christ-Had-Not-Died-For-Thee-Thou-Hadst-Been-Damned Barbon
I'd be very careful telling people what "rights" they have around their names.
EDIT: you can't really control what other people call you but you certainly can control what you call yourself.
Perhaps the actress from "Orlando" could start calling herself ~Swinton? Or maybe just "~"?
"Sorry, grandma, I know you've been sort of attached to your name for the last 80 years, but the white folks find it inconvenient for their computer systems. Don't worry, they promise they'll make something close for you."
Many of the clients of my ex-day job are married to legacy encodings like Shift-JIS precisely because they do think that their customers and students have a "right" to having their names written correctly. (Most of them also make a total hash out of foreigner's names, which I spent a good deal of time correcting. As far as I know my office probably still uses my name as test data, since it screws up about 80% of the systems we had, and it was cheaper to work around or patch than it was to fire me.)
The best solution is suggested in the other link: apologise for the technical limitations and offer a workaround. Don't demand that the world breaks solely to fit into your technical limitation.
If your name doesn't fit into Unicode, get a nickname. Bonus points if your nickname fits into 7-bit ASCII.
Say you want to order a package from somewhere. How does getting a nickname help? How do you explain it to the post office? I know it sounds like nitpicking, and it probably is, until it affects you personally. I've had my share of "name-mangling" and my name does fit into Unicode (not into ASCII though).
1. I don't want to inconvenience people with "weird" names.
2. I don't want to burden application programmers too much.
Requiring everybody to have a simple ASCII name would be convenient for programmers, but would be a big hassle for people whose names don't meet those requirements. "Be in the Basic Multilingual Plane or get a nickname" is a policy that, I think, provides a reasonable balance. Of course, supporting all of unicode isn't really that much harder, so I think that's a better balance.
I think the real bad assumption is that validation is necessary or even desirable - applying the technique brainlessly to names is the root cause of this problem - you don't really need to make any assumptions.
Not to mention systems that require usernames to have 6 characters or more, sigh...
The real question is: Do we stick to integers? Otherwise I want a supernatural number (http://en.wikipedia.org/wiki/Supernatural_numbers) or at least something like i (where i^2=-1), or perhaps a normal number (http://en.wikipedia.org/wiki/Normal_number).
Okay, from now on, all humans will not have names, they will be numbered in ascending order, starting from 1.
Boosh. If you want a 'handle' or 'username', you've gotta type it in whatever the encoding system supports.
Glad that problem has been solved.
"Who are you?"
"The new number two."
"Who is number one?"
"You are number six."
"I am not a number, I am a free man."
"HAHAHAHAHAHAHAHA."
Btw, why not use encoded pictures?