Falsehoods Programmers Believe About Names (2010)
kalzumeus.com
kalzumeus.com
So, now I drop the apostraphe.
In high school my school ID was completely broken. It printed <first half of last name> <first letter of last name> <middle name> <first name>.
They ended up putting a quotation mark instead of the apostraphe - somehow the machine handled that fine. That was the workaround.
Hotel wifi systems tend to break if they ask for a last name to connect. Again, training me to reserve the hotel under a 'fake' last name.
I actually notice it every time I write my last name. I'm like "is this a system that will break" and now I'm inclined to just never use the "real" spelling because then I'll get a mismatch later on.
This is something I've really noticed over the last decade. The older I get, and the more systems become a part of my life, the more systems I have to give a fake name. Paper records will show my true name, but I wonder if at this point, given digital records, what it would appear to be.
My identity is private. I use placeholder names for virtually everything I do.
Edit to be more substantive:
There are a lot of places where Know Your Customer rules apply. Further, there are places where identity really does matter, such as insurance. The approach you outlined sounds like a headache when it comes to those things. Not sure who the headache is for... :)
Why do you question whether identity would matter?
I can't just call up and say "My name is John Snow, I want to withdraw $100 from my account." Even a small system will quickly run into duplicates. Names are far from unique. Systems might use name for reference purposes, but beyond that, some other unique identifier needs to be used.
Over time all of the i18N people left for Apple. My login broke and I removed the accent. When I went to Apple I never bothered to add the accent again even though I was friends with the Unicode gang. I never added it back.
Which is the real problem - an apostrophe in a surname is a long, long way from unusual.
No technical limitation on the computer if it supports ascii. It’s bad programming and assumptions.
(HN readers, if you don't understand what's going on, then enable "showdead" in your profiles.)
"é" is not in ASCII, you're thinking of some "extended ASCII" variant, such as Latin-n (where n goes from 1 to 12). Those are insidious, as your "of course é is in ASCII" becomes "of course ř is in ASCII", or something similar. Finding the original encoding then requires out-of-band signalling, or guessing (heuristics ;)).
The technical limitations checked here are in bespoke code, not in the OS.
DDWRT WPA2 had issues with apostrophe in passwords. I wonder whether OEM firmwares have such issues, ironical considering special character does improve password strength.
O'Bryant -> O'Bryant -> O&#39;Bryant (screenshot: https://twitter.com/obryant666/status/1163638029212250112/ph...)
I'm also the person who gets irrationally annoyed at web forms that barf at my username+tag@example.com email address usage (+ signs being valid in the username part), so I can certainly relate to the ire, too.
This is really annoying because whenever I call about an account somewhere they will ask my name and I will say it is "AAAA ZZZZZZZ" and they will say, "oh it looks like I don't have an account for you". So then I will have to say, "oh wait, try 'BBBBB ZZZZZZ' or 'AAA BBB ZZZZZ', because that is probably me." This happens all the time when I book hotel rooms or reservations. I will come in and say I have a reservation under one name, they will say they can't find me, then i list several potential names. It causes all sorts of confusion and remarks of suspicion.
(those aren't my real names, just examples for illustrative purposes)
To make matters worse, my first name is broken into two words, plus then my family name. All three words can individually be perceived as "first names" or "last names". So the initial confusion, will frequently lead to mis-collecting of my name and mixing up the order of my name.
It is so bad, that I have had to register an alias now with the credit bureaus because it causes so many problems (yes you can do this if you call them) when authenticating me under various names.
Another fun/annoying recent story is that I am in the process of selling my house. My realitor told me that there was a problem because the county doesn't have me as the registered homeowner. My mortgage was signed using my real legal name. But apparently the county's computer system can't handle someone with punctuation in their legal name that is not part of a "title". So my name got dicombobulated in their system and is reported as a single letter with no punctuation. So the county records for my house do not show me as the homeowner, because their system made too many assumptions about naming conventions and my name got "broken" in the process. So now I have to correct it manually before I can sell my house.
Long story short... you should make no assumptions about names when creating a name field in a system architecture. It should literally allow for the absolute highest amount of flexibility possible. Especially if you are designing a system for say... the county government's title and deed records.
I'm Brazilian, a country where the vast majority of people will have at least two last names (typically one maternal, one paternal, but not uncommonly multiple from either or both parents), separated by spaces.
I've immigrated to the US, where the standard is for people to have a single last name, or if they have multiple, they're separated by hyphens.
To make my situation a little worse, legally I have two first names i.e. the second one is not registered as a middle name in any official document I have.
I've dealt with systems that wouldn't accept my last names. I either had to join them together, or separate them by a hyphen.
Some systems don't accept my legal double first name. When I got my driver's license in WA, the system accepted it but later had trouble generating the license identifier, so they had to manually apply the system's logic and enter it by hand.
Some systems just ask you to enter your full legal name, and then try to be smart about figuring out what your first, middle, and last names are. So I end up with the first of my last names being registered as a middle name, and I can't edit it.
Fortunately I'm close to being able to apply for naturalization, and that'll give me the opportunity to change my name. I'll just drop part of my first name and one of my last names and finally make it simple.
Edit: another thing that's funny wrt names in my life is my family situation. My wife has her own two last names (it's not super common in Brazil for women to take their husband's last name), and she has a daughter from her first marriage. So we all have totally different last names, with only my wife and my stepdaughter sharing one of their last names. That caused us a little bit of trouble crossing the border to Canada once.
Wherein someone named Amr keeps having his first name split into Mr. A by the airline booking system. Which would be bad enough, even if they didn't also insist that a name must be at least 2 characters. Good times.
See also the article about Christopher Null.
In school we had a guy whose last name was 'Van', a fairly common prefix for names in NL, which tends to be used like this 'van den Broek' or 'van der Berg'.
Teachers would ask for his name, he'd say 'Jan Van' and they'd invariably ask 'Jan van Wat?'. So he eventually gave up and answered 'Jan van Wat' right of the bat, which would leave the teachers even more confused because there was no such person on their lists of students...
Given that we know at least one guy whose first name is Van (Morrison) it is theoretically possible there is a person called Van Van somewhere.
At scale, its an easy option for helping users detect typos in their email domain (or for helping guide less technical users, like those who might provide address@gmail or similar instead of address@gmail.com)
(?:[a-z0-9!#$%&'+/=?^_`{|}~-]+(?:\.[a-z0-9!#$%&'+/=?^_`{|}~-]+)|"(?:[\x01-\x08\x0b\x0c\x0e-\x1f\x21\x23-\x5b\x5d-\x7f]|\\[\x01-\x09\x0b\x0c\x0e-\x7f])")@(?:(?:[a-z0-9](?:[a-z0-9-][a-z0-9])?\.)+[a-z0-9](?:[a-z0-9-][a-z0-9])?|\[(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?|[a-z0-9-]*[a-z0-9]:(?:[\x01-\x08\x0b\x0c\x0e-\x1f\x21-\x5a\x53-\x7f]|\\[\x01-\x09\x0b\x0c\x0e-\x7f])+)\])
Alternately, you can just use this super basic regex which handles all of the common use cases (while still allowing all of the weird edge cases), and functions as a quick "smoke test" for email addresses:
[^@^\s]+@[^@^\s]+
Basically, require at least one character (except a space, a line break, or an @) followed by an @, then at least one character (except a space, a line break, or an @).
The main flaw this simple regex is that it is too permissive and potentially allows weird edge cases like special UTF-8 characters that might not be allowed.
But hey, one might argue that if a user is intentionally entering an invalid email address with wonky characters to create an account, then that's kind of their fault when they never receive the account activation email with the email validation link, right?
And if you are going to do that email validation step anyways, you could probably get away with just using the simple highly permissible "smoke test" regex for your email address input form validation.
"f@@bar"@example.org
but it's not accepted by the basic regexp.The first regexp doesn't look correct to me either.
[^\s]+@[^\s]+
I guess you could get even more crazy and go ahead just do the most bare-bones check possible really: .+@.+
And you were right. The first regex got screwed up when I copied and pasted it. Here's a better source from https://emailregex.com (?:[a-z0-9!#$%&'*+/=?^_`{|}~-]+(?:\.[a-z0-9!#$%&'*+/=?^_`{|}~-]+)*|"(?:[\x01-\x08\x0b\x0c\x0e-\x1f\x21\x23-\x5b\x5d-\x7f]|\\[\x01-\x09\x0b\x0c\x0e-\x7f])*")@(?:(?:[a-z0-9](?:[a-z0-9-]*[a-z0-9])?\.)+[a-z0-9](?:[a-z0-9-]*[a-z0-9])?|\[(?:(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?)\.){3}(?:25[0-5]|2[0-4][0-9]|[01]?[0-9][0-9]?|[a-z0-9-]*[a-z0-9]:(?:[\x01-\x08\x0b\x0c\x0e-\x1f\x21-\x5a\x53-\x7f]|\\[\x01-\x09\x0b\x0c\x0e-\x7f])+)\])
I suppose there isn't a whole lot of value even using a quick "smoke test" regex, when the big version is already a solved problem, and can be so easily copied and pasted at this point...If you are lucky the regex will just cause maintenance issues. If you are unlucky it will stop users signing up.
n. Names can be included in any sentence in the same form as user entered them.
For example any language like Russian where names decline breaks it.
n+1. Names are grammatically nouns
In Russian last names and patronimic/matronimic names are usually adjectives
n+2. Any person's name can be disclosed to people who know them by their other name.
For example if some LGBTQIA person has homo/bi/trans+-phobic relatives they financially depend on then such disclosure can be a catastrophic outing.
When dealing with persons, it's necessary not to confuse their roles.
I'm pretty sure most US legal documents required some form of given name and some form of family name. If you're trying to interface with as US government system, and that's your primary use case, a single first and last name will probably suffice (assuming it's not too arbitrarily short, people can change their name, etc.)
My birth certificate has all of these as well as my passport. My drivers license doesn’t have the é.
Truncating my name is how I get mail from people states away who have never lived here. One of them being my dad. He’s never lived in this state so it’s not due to an old address.
That’s also how I commit felonies by opening those letters thinking their for me.
While delivery and the nursery apparently are set up to deal with this, (by barcoding the babies patient bracelet) computers in other departments were not, at least without navigating multiple menus and screens.
Every computer that required patient name and birthdate to be entered was a 10 minute process, to the point that we were famous by the third day of visits.
At some point, a higher up was able to log in and "fix" the name to be discoverable by just last name and birthday, or via mothers patient info. I'm not exactly sure how, but I was impressed with it being "fixable" so easily (for us at least), especially network wide.
I also understand now why the vital records people were so persistent in trying to get us to finish the birth certificate before discharge. We obviously didn't.
Joseph Smith // nice and generic
Joseph Smith Jr. // gotta make sure we tell people he's the son of another Joseph Smith
Joseph Smith Jr // psh, we don't use periods!
Jose Smith // if from a different country, some sites keep their native name
José Smith // yup, sometimes accents are kept
Joe Smith // nicknames, all about those.
Joey Smith // kid nickname, sure
Jos� Smith // literally even had to deal with a site that had unknown unicode letters
Juan Jose Smith // sometimes, people go by their middle names, but some sites like to use their full name instead
Juan Smith // remember how some people go by their middle names? Maybe a site only gives you their first and last
How, how in the world is someone supposed to go through and match all those names with each other in a somewhat quick manner? I can't imagine it's fully solved because there'll always be other cases.I've tried a couple ways, like removing periods and accents, removing "Jr", "Sr", "III", removing all vowels even. And then went further and tried some of the string matching libraries that return a number with how close words are to each other. That would help cases where Joe, Joey, Joseph would come back quite close, but then I'd run into the issue of a "Darren Smith" and "Darek Smith" would come back so similar the computer thought they were the same person.
In cases like names, that Patrick writes about here, yeah, they're almost impossible to get right and it ends up being mostly by hand, which is fine since I can get it correct, but eats up so much time.
And this is from an industry with guilds and unions to try to ensure that no two credits have the same written name.
You can make a very good guess, but names can't tell you enough to determine if these two people are the same or different. One person can have many names, and many people can have the same name.
A very efficient way of matching strings across typos (aka "fuzzy") is SymSpell, but just finds nearby lexical matches efficiently.
Sorry, couldn't help myself.
Names are a reasonable indicator of the person's sex. Go ahead and prefix a Mr. / Ms.
My fav example of this is U Thant, who was Secretary General of the United Nations.
His name was actually just “Thant”. The U is something akin to “Mr” in Burmese.
I cant imagine how often he must encounter systems that get tripped up by that.
In the web sites I work on, I have decided to have one field that asks "How would you like to be called", and allow full Unicode range in it. If you want to use Emoji, fine, go for it. I have fewer data to store (privacy and data protection concerns), and users feel better not typing the last name either.
(And then of course you get The Knights Who...Until Recently Used To Say Ni)
https://egov.ice.gov/sevishelp/schooluser/machine-readable_p...
Actually, names should be considered to be descriptor data that can be searched on, not as identifiers.
http://randomtechnicalstuff.blogspot.com/2014/03/google-name...
What about an infant before they are named?
In this case, it would be quite normal to have an 18 month old with no name at all. I mean, you could call them "Baby Surname" or something. But they don't have a name.
All MRNs should be unique and everything else can be a duplicate.
Every system should support name changes and merges. John Does come in all the time, and then you figure out their existing MRN, or their real name.
> Every system should support name changes and merges. John Does come in all the time, and then you figure out their existing MRN, or their real name.
In theory yes, I’d hope to see that across the board, but fat fingering happens way too often on top of the reuse of mrn’s. I work in the IHE space, so I can point fingers at most of the big guys as we accept their HL7/CDA feeds and not even ids from the same system are consistent between the formats. I don’t want to get too off topic with that last line, but yeah we try to consolidate records that are received by leveraging Fellegi-Sunter[0].
https://shinesolutions.com/2018/01/08/falsehoods-programmers...
Well, what hope is there then?
No it's not.
Ok I'll bite:
A defined amount of space: 250 Terabytes.
Please give me an example of an extant name which does not fit into that amount of space.
Bulletproof edit: or the physical space that is the facade of the empire state building.
That was absolutely insane. Try applying such logic in China or a big Arab state.
The "birthday paradox" means you need just 23 people before you have a 50% chance of a collision, and 70 gets you 99.9%...
Uh.. can't help you with that one!
The point being that 'name' is a term interpreted in culturally varied ways... which when I design most systems I dont care about. so to make it clear, lets not use that any more.
Any examples where people don't have names?
Some tribes also don't assign names until children are nearing adulthood, calling them what basically amounts to child/boy/girl.
Other similar examples: https://news.ycombinator.com/item?id=20676904
How is this not true?
Edit: it's a joke lol, as I get farther in the list
As far as I can tell, all of those are not jokes.
[0] https://archive.nytimes.com/www.nytimes.com/2009/04/21/world...