Until then, can we agree that getting things right for 99.999% of people in your target market is good enough, and that, although unfortunate for people with culturally different names, the time and cost of getting all 40 points right would be too damn high?
1) Names are context-sensitive, so ask for a person's name by giving the context you intend to use it in. Something like, "How should we address you in an email?" or "How should we display your name to other people on the site?".
2) Do not apply any sort of transformation to a name. Don't parse it, don't truncate it, don't lowercase it. Treat it like an opaque string and use it verbatim.
3) Databases do need limits on the length of a column (unless you use a blob...) so be generous with the width of the columns. The longest recorded name is 746 letters long (http://en.wikipedia.org/wiki/Wolfe%2B585,_Senior), so make sure it's longer than that.
4) I think the no-name case (#40) and the not-in-unicode case (#14) are fairly obscure, but if you want to handle them then then make names optional but require an avatar if they do not provide one.
5) Allow a person to change their name at any time.
I don't know if it's still the case, but Facebook used to have this annoying bug where they'd not only capitalise the first letter of my surname (incorrect!), they would then not allow for me to correct it because it seemed to have been treated case insensitively. They don't seem to enforce the capitalisation now since I've created a new account though. I know it was just a bug, but it was so disrespectful somehow.
That's not very hard, and is in fact easier than trying to validate or restrict names. You don't have to learn about all the different cultures etc, just don't assume they are like yours.
And 2048-character limit on that field ought to be enough for everyone, right?
But that's exactly the issue here; there aren't many people with names as long as Janice's, but when these people have problems with computer systems, we get to hear about it.
Why do you need a fixed character limit at all?
Each row contains the same columns. For example: table Person: col1 label: name / type: text / size: 10 2 label: age / type: int / size: 3 3 label: address / type: text / size: 200
the type tells the database how to encode the data found in that column - an int takes less space than a string, so this is a space saving measure. The size refers to how many characters are allowed in that field, because that is how much space will be allocated when the row is created. If you don't need 200 characters of space, it's wasteful to allow it. But you kind of have to put a number in there, even if it's a big one.
Correct me if I'm wrong, all my knowledge comes from Dr. Marla Weston's 2nd year Databases Design course.
http://www.postgresql.org/docs/9.3/interactive/datatype-char...
A character with no representation in the latest Unicode standard.
Learn to fix the problems you can fix, and only the problems you can fix. Ignore the problems you cannot fix because, as per hypothesis, you cannot do anything about them.
In other words, what if you get hit by a micrometeorite on your way to the bathroom tomorrow?
Sorting is another area where it'll vary a lot depending on locale (perhaps you want to sort on last name for example, or not depending on your users) - it might be best to try to guess that from whatever they enter, rather than imposing a structure on them.
Personally I'd choose the simplest solution which works, so no last name/first name, just a name field for display, depending on the requirements, a website/app might need to collect more of course, but for most websites there's no reason to try to control user names or split them up into arbitrary first/last pairs.
More on topic, why should official systems allow for names of arbitrary length? Obviously your 1800 character preference for last name is not going to fit on all standard forms of identification. A standards-compliant name should be issued and you can use your preferred name for personal matters.
I cannot adequately express my dislike for this kind of mockery. Seriously, fucking stop it.
Scribble: to cover with scribbles, doodles, or meaningless marks. Names aren't meaningless, and TMK names don't change depending on the time of month. By phrasing things the way you are, you are making generalized (and fictitious) blaming statements about something that is very personal. And you got called out on it.
Addressing all of the issues on the list pretty much boils down to accepting a scribble as a valid name. And if you are going to accept point 8 on the list that people's names change at any non enumerable event, then that is just as difficult as a name that changes based on the phase of the moon. Worse, actually. Moon phases are at least predictable. The point is, I'm creating a plausible scenario that fits the suggestions in the list. It is not necessarily fictitious, and it is absolutely in no way worse than the list's expectations.
Please identify how I made any type of "blaming statement". If you're going to call me out then go ahead and call me out. As far as I'm concerned I've been accused, not called out.
The real issue is that too many programmers apply their own cultural norms to the issue of naming, in ways which privilege their particular culture at the expense of others'.
(You want to be able to identify people uniquely? Fine: allocate them a globally unique ID string within your own system. Then let them link this to whatever the hell they want -- be it an Anglophone forename-middlename-surname string sequence, or something entirely different that fits their cultural requirements. Once you get into designing a procrustean name storage architecture and insisting that everyone uses their true name, then you are inflicting your own culture's naming requirements on them, which is offensive at best. Especially when it isn't necessary.)
Most systems aren't going to be built with infinite flexibility. Some might argue that wouldn't even be possible. And perhaps it has absolutely nothing to do with the programmer's own cultural norms, but the system they are programming on, the budget they are dealing with, and their level of competency. Not every action that is offensive to a particular group is done deliberately to offend that group or even with the knowledge that it could be offensive.
Irrelevant.
>or even with the knowledge that it could be offensive.
Irrelevant to your argument, which is that we should continue doing it even after we figure out that it's offensive, because fuck them.
The system was written to work for the 99.9% of users it serves, the other 0.01% will unfortunately have to deal with the problem of being outliers.
Where's that quote about lack of empathy masquerading as "brutal honesty/pragmatism/anti-PC-ness"...
What a relief that nobody is advocating infinite flexibility, but just pointing out where specific lack of flexibility is problematic.
There's no such thing as a free lunch. Optimize for the use cases you expect: if you're making a phone book for middle America, that's a different problem than a social network for Russia.
I would also venture to guess that not having a SSN or even an ITIN to use in place of the SNN is an automatic disqualification for attending school, getting a driver license, buying a house, or getting a job.
> People’s names are all mapped in Unicode code points.
At this point we've gotten into absurdity. Yes, supporting Chinese people and Hawaiian people with superlong names and people with first/last vs last/first or hyphenated or non-alphanumeric or whatever... all the stuff that can be supported with a single string of UTF-8? That's all reasonable, and there's a very real progrem with programs that don't support those kinds of names.
But going over the line into names that can't be expressed in any form of text known to a computer? At this point, the ludicrous exaggeration is only barely further than the silliness included in the "Falsehoods" list.
He's not mocking "you have to support a half-German half-Chinese person who was raised with an Inuit nickname". I'll happily support Prof. Dr. Ms. 陳 (ᒥᓂᔅᑕ) Wolfeschlegelsteinhausenberdorft III. That's fine.
But the "Falsehoods" article says that isn't enough, and is pretty much describing "seven different names, each consisting of some kind of personal scribble that changes depending on the phase of the moon".
Normally I'm the first to complain about racism and those kinds of insults, talking down to people who are just too much trouble to bother with. But that "scribbles" comment is literally outlining the lengths the article describes.
Really? Because a space is commonly used as the delimiter between first, middle and last names. Wouldn't you think that, at a minimum, there would be some ambiguity about whether Jones was (one of) her middle name(s) or the first part of her last name?
Seems easy enough, right? They don't have one big field for [ANN BETH SMITH JONES] and then try to guess what name is forename or surname.
Some systems don't like hyphens, or don't like spaces, or appear to accept them but silently drop them kludging the two names in SMITHJONES.
Add the 'scunthorpe problem' and things get even more frustrating.
One of the Google feedback pages declined my feedback because my (legitimate, linked to my Googlemail account) email address and name contains "Cocks".
http://www.kalzumeus.com/2010/06/17/falsehoods-programmers-b...
de la Chapelle
von Braun
zu Guttenberg
af Slestad
A computer system that assumes that Wernher von Braun's middle name was "von" is ridiculous.
Of course, apart from using the proper name/title when composing a letter (something like Freiherr von Braun), which is easy if you know the ‘last name’, the computer system in question would also have to sort this guy properly next to Werner Braun.
Names are complicated.
Take John von Neumann for example.
* Born 'Neumann János Lajos' (family name followed by given name order used in Hungary).
* His name changed to 'margittai Neumann János' when his father was elevated to nobility.
* He germanized his name to 'Johann Ludwig Neumann von Margitta'
* ... which got shortened to 'Johann von Neumann'
* He finally anglicized it to 'John von Neumann' when he moved to the US
I'm not at all sure I got that 100% correct...
The following mostly-identical names read very differently:
- Sarah Anne Jones Smith - Sarah Anne Jones-Smith - Sarah Jones Smith - Sarah Jones-Smith
Or how about Joran van der Sloot?
My example was meant to be "the final word and all preceeding lowercase words are considered the last name."
Also from French Canadian, some last names have a period in them, I know someone with the last name "St.-Germain".
(PS Not our actual names, but they follow that pattern).
That point being that I was surprised that it wouldn't even occur to someone that there would be some difficulties with this in the world in which we currently live.
Blaming the user whose name is too long for that field rather than the mistaken assumption in the design of the system is not.
Where? Everywhere in the world?
Anyway, I thought that the reasoning that a customer could not be served because address or name "does not fit into the system" were frowned upon since the time of COBOL mainframes, but seeing this whole thread with its resort to a hyperbole of "0.01%" cases is a welcome contrast to the otherwise high aiming goals plastering the homepage of this site.
1. Yes, this is sometimes known as discrimination.
2. Yes, you have no way of knowing the size of the market if you are excluding them - this may be disguising a group that's significantly larger than 0.01%.
3. Yes, if you cater to these people when your competitors cannot/do not, you have a captive market.
4. Yes, everyone is 0.01% in some fashion - if you continue to apply this reasoning, you will eventually exclude everyone other than yourself.
2. If you want to argue for an arbitrary budget increase to satisfy an unknown market size that may be larger than 0.01% whose requirements are drastically different than the others - go for it. I'm not going to do that.
3. Congratulations, you win the possibly larger than 0.01% market that I ignored. In the meantime I've iterated on several other useful bits that the vast majority of users care about.
4. The number of input fields in my dataset isn't infinite.
And bureaucracy always ends up annoying people.
Obviously you want to aim to be flexible, but I don't think there should be an obligation, for example, to support characters that aren't in unicode.
> A standards-compliant name should be issued
>More on topic, why should official systems allow for names of arbitrary length?
These are really incendiary statements (a local may accuse you of being insensitive[1]) that places the fault of the issue on the aboriginals of Hawai'i for observing their traditions rather than the programmers who should have known better. From what I understand DB's get re factored all the time and if the drivers license can fit her name I fail to see the issue.
[1]euphemism for racist
You just accepted a length limit. Congrats on being culturally insensitive.
If we can't expect the name to be in a character set, why can we expect it to even be two dimensional? What if it's an object? What if it's a sound that Western voices can't make without years of training, ala Tuvan throat singing, with no written form? What if it's a sound that needs to be played on an instrument known only to the person's tribe/nation/whatever, and the only one in existence is a family heirloom halfway around the world? What if it's something like a particular pattern of brainwaves created by meditating a certain way? Maybe a little extreme. There's an African language which is primarily location-based. People are raised from birth with an extremely accurate innate sense of spatial position through dead reckoning. Greeting someone and identifying yourself requires specifying your location and the direction you're facing. Fascinating stuff. There's a great Radiolab segment about it: http://www.radiolab.org/2011/jan/25/birds-eye-view/ So is it racist to choose a column type for holding names that isn't designed for GIS? If you did, isn't it racist to use a system based on Cartesian coordinates instead of the native way of representing locations?
My point is that you cannot possibly cover all the edge cases for the ways people identify themselves, and calling that offensive is just silly. Giving people identifiers that you created is also offensive - that's prison/Nazi crap. So what do you expect us to do?
A blank box in the character set of the primary language(s) where the application is used with zero validation and zero transformations is a reasonable demand. Calling it racist and offensive to fail to cover every edge case is a bit of a stretch.
Accommodating everybody is a pretty tall order (see: Prince). Where's the appropriate place to draw the line?
Half the times it's "falsehoods a programmer back in the '80s believed about names and I have to operate with that goddamned system".
That and we have to support all the cavemen that want to sort by last-name. I'm perfectly happy to just use one 1000-length UTF-8 field for names and skip all the stupid abbreviations and worrying about firstname/lastname, but for some reason the BAs insist on supporting "Mr. So and so" and being able to write "Lasty, Firsty" as if that is in any way useful when you've got a searchable document.
I suppose you might occasionally need the Lasty, Firsty when printing into dead-tree-form, but when have programmers ever cared about that?
Sometimes when you live within a culture you will need to adjust to the way they do things.
You should have paid more attention in CS school!
You can alway store name-elements as meta-fields as long as the original name is untouched.