Some problems with 'first name' and 'last name' fields in data
utcc.utoronto.ca
utcc.utoronto.ca
Collect two values:
1. What do you want to be called? e.g. if we send you an email we'll say Hello {name}
2. What is your legally recognised name i.e. what's on your passport and official documents.
#2 is only relevant in some domains, in many cases #1 alone is sufficient.
1) People can have multiple passports with different names.
2) a single passport can contain multiple names for the same person.
AFAIK, the MRZ and chip don’t support alias fields?
3. A string that will never be displayed, but will result in proper sorting. It's not simply a concatenation of all name parts, though. Omitting spaces and punctuation is also critical for a nice sort, so that apostrophes and the like don't shunt things to the top or bottom. But don't just filter to Latin letters, keep all [0] letters!
When my last company did phone numbers we just let it be free text…
I've even run into trouble occasionally as a Western person with fairly normal first and last but no middle name. More than once I've been confronted with online forms that require a middle name (not for many years though, thankfully).
Good times.
There is an ISO standard for phone numbers, actually. There are even libraries to parse phone numbers into the format. That being said, for “non machine readable” phone numbers (say for a contacts app) some say you should keep it free text because people plug all kinds of shit into the field and if you aren’t using it to send messages/voice/whatever just let the user do whatever…
The problem arises when the order is reversed (e.g Vietnamese), or when there's only one name and the form requires at least two names. And probably for other name formats saying around the world that I'm not aware of.
First: Miquel Middle: empty Last: Martí i Pol
The problem is for names that don't work in that format.
E.g. First part of my name is Muhammad but no one ever calls me by this in my country (Pakistan). This is one of the most common prefix here. I am usually surprised when some one here in UK correctly calls by my name instead of Muhammad.
Filling official form with Muhammad as my "First Name" is a weird experience knowing that this is what I will be called in communications.
If someone tells me their name, I think that it is right to believe them unconditionally. No matter if it starts with Wang (100M people!) or ends with -son/-dottir (~all of Iceland) or contains affixes or other common parts. Names are deeply significant and have more value than just contributing information entropy to a unique identifier.
the icelandic case is different, as it does after all function as a lastname, just a suffix on the lastname.
I see no value in this, its kinda like putting salutation into the name. I would receive NO value in having an additional firstname of "Mr"
Names are part of identity the same way origin is. Can't just ask to purge it because everyone around has some of it in common.
this is why someone might have been called "johnny 1 leg", because he only had one leg. You do not hear things like "2 legs 2 arms 10 fingers 1 heart 1 bellybutton johnny"
Then you can have, say, "Western" naming convention, "Asian" (reverse order), "Spanish", "Catalan", "Hungarian", "Ancient Roman", and some of these are going to look very similar. And then if Madonna Ciccone or Beyonce Knowles comes around, you just use the catch-all for the mononymous.
It's not an elegant solution and it's not going to work for user-oriented forms, but perhaps on the backend, I don't know, for Wikipedia or something. Wikipedia already tags many biographies with the naming convention that is used.
Names are effectively a collection of arbitrary Unicode values with a fair number of complications (sorting, multiple names, etc), and there is no set of standards to draw upon for implementation.
Fair point. I used those examples mostly to acknowledge that programming has moved on from treating them as strings and integers. It's true that in the US there are few, if any, restrictions on naming. Only in 2017 did California, by law, allow names with diacriticals, as in José. However, in other countries there are interesting restrictions and patterns. Icelandic law, to give one example.
> Names are effectively a collection of arbitrary Unicode values
Well, yes, for just the string, but names are more than a sequence of characters. The structure of what a name is, and how it's presented, matters. for example given vs surname order, patronymic, Mongolian clan names, supported characters in CJK, and so forth. Those are things that are hard to get right, and having every organization doing i18n create its own one-off name handling code has led to chaos and a lack of interoperability internationally.
- Every programming language, database, etc implements something, but overall they’re not the same solution, leaving us mostly where we are today.
- Some computer standard emerges, but fails because it’s not flexible enough for all the business and cultural needs, leaving us mostly where we are today.
- Some computer standard emerges, is good enough for most needs, but leaves people bitter because it leads to yet more cultural standardization in a world already suffering from cultural globalization.
I’m inherently skeptical of all change, I suppose, and my belief is we’ll just muddle along forever.
Maybe my experiences in the industry have been particularly plagued by the handling of names, sometimes so badly that I have to roll my eyes. The handling of human names just seems like a big opportunity to remove friction across international boundaries.
For example, Ellis Island was a huge source of novel spellings, loss of diacritics, Anglicizations, and other rather atrocious word crimes. However, since perhaps a small, finite number of civil servants were the ones transcribing the names and putting them on official US documentation, this had the effect of basically standardizing a lot of names for families now living in these United States, and possibly reducing the variation that was found among them. While I haven't directly seen many Ellis Island records, I've seen a lot of US Census records, and it's obvious that the census workers always wrote down sort of what they heard, and when it comes to names, there could be all kinds of phonetic trickery going on.
This still goes on today, and anywhere you have a bored, underpaid civil servant entering names into a rigid database system, you'll have a rather unpleasant standardization on the lowest common denominator. People who speak Arabic or Indic languages may often choose, on-the-spot, how to romanize their names. There are some romanization standards, of course, for most popular scripts, but why not follow multiple ones in the same word?
I imagine that it may go the other way, as well. Whatever country people are immigrating to, whatever written language that database is implemented in, names will get standardized to that lowest-common-denominator.
Focusing on the characters is being stuck in the name-as-string: the exact problem I'm advocating we give up.
I ran into some stupid company policy where I had to be identified by first and last name...
Resulting in people thinking my name is "son".
The I had to explain that no, son is not a name, nor a surname. It is an adjective, not a substantive, it's sole purpose is differentiate me from my father, when needed. Calling me just "son" makes no sense.
I pointed out to that company that if they applied that policy to bill gates, that is names Willlam Gates III, they would end Calling him number three, and it is obvious that number three is not his name.
Then there is the fact my "middle" name that matters. My first name was just to honor a famous catholic priest and is not even a valid word in my native language. Nobody calls me by my first name... unless there is some bureaucracy somewhere, then they insist on calling me by my first name, and son as my surname.
1. The "Y2Gay" problem
Was a problem because databases were designed with the idea that marriage could only be between a male and a female.
2. Falsehoods developers believe about names
From HN's legend: patio11
https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
When it comes to these considerations for instance, it's possible to have a first name, but no last name and vice-versa.
Also, never mind how many words make up a first or last name, but what's the minimum number of characters allowed for each?
What characters do they allow, and which shouldn't they?
Also, what version of the person are they attached to? Is Dr. J. Stevenson the same identity as Jonathan Stevenson, or Jon Adam Stevenson?
Once you have the data, how you present it as a separate i18n thing. Do you want to address an <ethnicity> crowd by their last name? Cool! That’s a rendering issue, not a database problem.
Though I do agree that it's mostly a matter of i18n. In Spanish you should ask for "nombre" (name) and "apellido(s)" (surname), if you say "given name" and "family name" users will be confused.
Preventing people from writing their own names because you decide to use a regexp from stackoverflow is also bad.
Previous to that I worked at place with a key customer in Japan, so everything was pretty well i18n and I got used to seeing given name/family name or similar. I discovered I had muscle memory for typing 'surname' instead of 'lastname'.
Although the proposal still doesn't solve all edge cases. And the current edge cases have presumably found a solution, like my girlfriend that has 2 last names, and is from a 'western' country.
My parents address is stupidly long and doesn't fit in to most address boxes, so you also need to allow arbitrarily long addresses.
Oh, and your email validator is wrong too.
To understand this as an English speaker. Imagine people aren't called John Smith, but rather John the Smith and not Bill Boston but Bill of Boston. Now imagine that a large percentage of people had an infix in the name. It would be impractical to order stuff by last name and find a massive index on the letter "o" because all the last names starting with "of" or at "t" starting with "the". To balance this out, parts that would translate to words like "from", "the", "of", "to" or "on" would be considered an index. Part of the last name, but not considered part of the name index. John the Smith would be found at the letter "S" usually structured as "Smith, John the".
It's the iTunes "The ..." problem.
Thanks for clarifying.
Now, I'll admit: it actually did thwart my attempt at not providing my actual email address, for a good several attempts. I had to increase my creativity to a level I truly didn't anticipate, before finding something it accepted. And I definitely had enough randomness that it wasn't just disguising an "already exists" error with an "invalid address" error for privacy purposes, I am certain.
What in the world kind of validation is this? Do people really end up resorting to entering their actual email address often enough to make this at all useful?
I'm constantly annoyed when asked for both my legal name and my preferred name that companies choose not to address my by my preferred name. That's why I gave them my preferred name, after all.
"Your immediate supervisor's name must only contain letters (A-Z), hyphens, or spaces."
(At least it didn't actually impose the apparent restriction that the name must be ALL CAPS.)
Just store a name field, and if you need to interface with other systems with multiple name fields, ask the user to fill those fields out.
Even when people adhere to that norm rigorously (which they don't), it's completely unusable as a unique identifier, and is of dubious utility to disambiguate two distinct humans in many scenarios?
That there's not even any real agreement on spelling (especially as, with some, it will need to be spelled in multiple alphabets)?
If this stuff's not old news to you, you haven't ever worked with it.
The one that will really make you want to curl up in a corner and whimper are the problems with birthdays.