My last name makes me invisible to Computers (2015)
wired.com
wired.com
If web forms aren't accepting NULL, then somebody probably specifically programmed the word 'null' into a filter of disallowed entries. Probably to stop clerks from entering the word 'null' to mean empty string. This has nothing to do with null being a reserved word in many languages, I'll bet the forms that aren't accepting 'null' aren't accepting 'none' or 'empty' either.
Is this ActionScript that's doing the weird xml tag interpolation?
I stand corrected, this is definitely possible. Still, I bet this person's real problems stem from explicit coding.
var nullXML:XML = <root>null</root>;
if (nullXML == null) { trace("Some XML is null"); }
nullXML == null is true!
That's... impressive.Edit: It's true in PHP though. :-/
Can you post exactly what you did in PHP that shows "NULL" is equal to NULL? There's quite a few approaches like ==, ===, and is_null(). I can't get any to think "NULL" is NULL, though I imagine I'm missing something.
Unfortunately it's replaced by something even more messed up
Object a = null;
String s = "foo bar " + a; // s == "foo bar null"
So, it's possible to get "null" as a string downstream, when some variable that should be non-nulleable - was null. If you find such a bug and incompetently fix it by checking for "null" downstream instead of checking before turning variables into strings - you have the error from the article. Especially if you check ignoring the case.BTW guess how I know it's working like that :)
So I guess if you have a comparison that somehow casts null to a string before comparing, you could run into this issue, but that's still bad programming.
I used to own null@myundergrad.edu as an email alias (with my real university, not that generic .edu, of course) and I got all sorts of interesting things... but that was very intentional on my part.
See:
"null".equals((String)null) //false
"null".equals((String)null+"") //true
System.out.print((String)null) // throws NPE
System.out.println((String)null) // prints "null", I guess because it appends "\n" insideBTW there's another "fun" gotcha, when you interface java code and oracle database which consider empty string and null to be the same thing. Depending on how you handle data from database you end up with null, "", or "null" :)
And indeed,
public static String valueOf(Object obj) {
return (obj == null) ? "null" : obj.toString();
}
http://hg.openjdk.java.net/jdk8/jdk8/jdk/file/687fd7c7986d/s... System.out.print((String)null)
Does not throw NPE.http://www.javarepl.com/term.html
But you're right, in regular java environment it doesn't now that I've checked.
I've known Java to be a bit verbose, but I thought it was mostly a reasonable language aside from that. I might have expected something like that from Javascript or PHP. Never in my wildest dreams would I have thought that boring old Java would interpret a null cast to a string as a literal string "null" when combining strings. Even Ruby and Python don't do that - they throw type mismatch errors instead. C# treats it as empty string.
In Python:
"abc" + None
TypeError: cannot concatenate 'str' and 'NoneType' objects
"abc".join(None)
TypeError: can only join an iterable
In Ruby "abc" + nil
TypeError: no implicit conversion of nil into String
In Lua = "abc" .. Nil
attempt to concatenate global 'Nil' (a nil value)
But not Java. Java implicitly tries to coerce the provided value to be a string, which seems out of place in a language that values type safety. That all types might also be null also seems out of place in a language with strong, static typing, but that's been discussed to death. The problem is compounded by the fact that null is actually converted to "null" instead of an empty string. bar = "bar"
foo = "foo"
foo_bar = foo + baar # intentional typo (baar will be None)
In Java this would be a compiler error, but python accepts it until the program is being run.My claim isn't that Python is safer overall than Java. Instead, it's that Java, a language that is mostly type safe, most of the time should not have these two potentially surprising behaviors:
1. The standard string concatenation operator does implicit coercion rather than rejecting an input that isn't a string. There should be a builtin to make a string from any value no matter what for logging and debugging, but that shouldn't be the standard concatenation operator.
2. This is more controversial, but strongly statically typed languages should not allow arbitrary values to be null. That sabotages one of the major strengths of strong static typing. Instead, there should be an option type to make it explicit. For something familiar to most programmers, SQL does this.
It's like if a guy you were friends with since college and always knew to be solid and reliable but a little boring, and now many years later, he's happily married with kids, watching sports, working in a large hierarchical organization, doing pretty average stuff but nothing unusual or exciting. If you suddenly found out that guy had secretly been a Furry the whole time, that would be really shocking.
Not that there's anything wrong with being boring and reliable, or being a furry. But the sudden change in how you saw somebody or something is stunning.
I half expect anything written in Java to be littered with AbstractFactoryFactories and giant frameworks using 20 different design patterns to write Hello World. I never would have expected Java to silently convert an actual null to the string "null". I always thought the one thing you can count on Java for was to be strictly strongly typed, and never silently do any weird random conversions that nobody would have expected. Guess I was wrong.
Though poking around in a few other languages, JavaScript does indeed do that also, though I kinda expect JavaScript to do things like that. In Ruby, nil.to_s gives empty string, and adding a nil to a string gives you a type error. In Python, trying to add None to a string also gives you a type error, but str(None) does give you 'None'. That's a bit disappointing, but not shocking to me.
Javascript does it too, but PHP doesn't.
I don't see the problem with that. Strings that contain formatted variables should be used for display only. Besides, when someone inputs his name, the input is already a string, and I don't see how you would dereference string contents.
The problem is downstream, where some galactic idiot has special-cased the string "null" to turn back into an actual `null`.
>"hey "+null
"hey null"
I wonder why the language designers thought that was a good idea?For some reason someone decided to check for null instead of the stricter option of throwing a NullPointerException. Maybe to help debugging or to follow the gist of toString().
I guess it was too late to change even in Java 0.9... Autoboxing, iterators and other newer features are more strict,so I suppose it's just one of the few irregularities left from prehistoric times...
> git grep -i '"NULL"' src |wc -l
> 32 First,Middle,Last
Alice,Q,Foo
Bob,NULL,Bar
In a world like that, some code got written in the wrong tier to output value.toUpper() == "NULL" ? "" : value
and you wind up with a web form that can't roundtrip a name that contains the word 'null'.It's one of the reasons I wrote a pared down 'dumb' YAML parser that assumes scalar values are strings unless directed otherwise: https://github.com/crdoconnor/strictyaml
(I just received some files like this. One is a file that has names. If someone had a name of NULL, it would go into the database as a null.)
This post is a classic on various name issues: http://www.kalzumeus.com/2010/06/17/falsehoods-programmers-b...
Manpages aren't written in TeX though (apparently they use something called "roff"), but they also contain things like `read' instead of 'read'... perhaps manpage writers tended to like TeX too?
http://www.read.seas.harvard.edu/~kohler/class/aosref/ritchi...
"Programmers use the grave accent symbol as a separate character (i.e., not combined with any letter) for a number of tasks. In this role, it is known as a backquote or backtick."
paul@tal:~$ od --format=x1z tmp/tonos-oxia
0000000 74 6f 6e 6f 73 3a 20 ce ae 0a 6f 78 69 61 3a 20 >tonos: ...oxia: <
0000020 20 e1 bd b5 0a > ....<
It's super fun when you are the only tech person in a Classics grad program and everyone else is turning in papers that look like ransom notes, because every fourth vowel has a different x-height from all the other letters. :-)Here are the two Unicode ranges (PDF):
modern Greek: http://unicode.org/charts/PDF/U0370.pdf
ancient Greek: http://unicode.org/charts/PDF/U1F00.pdf
Then again, try explaining to tech people the difference between ή and ἠ and how you can get ἤ or ᾔ. :-)
How do you enter them nowadays by the way? In the early days of the internet before unicode there were special fonts from SIL for example that first of all were using Latin characters (you'd type W and it would look like Ω) and secondly I think had the diacritics as separate characters, so the font would combine them. It was messy but at least you could type it on a normal keyboard. You could even read it in its ASCII form, most of the hacks were pretty reasonable.
Personally I type Greek using a vim keymap file, usually when writing LaTeX. I believe my keymap file is influenced by those older non-Unicode fonts, because I type w for ω and ;h for ή and >~h| for ᾖ. But my fellow students would mostly use Word. I don't know how they would type the letters, but commonly they would mix the tonos letters with everything else. I think what was happening is that Word would automatically substitute fonts that offered those codepoints, so it would end up showing Times for the letters with tonos and Palatino for the others (or something like that). Hence the ransom note effect.
Source: have a hyphenated last name. "No special characters in this field".
Ups and FedEx charge shippers ~$15 for each instance where they have to address correct.
So, yeah, the period thing is dumb, but automated correction is hard. Even experts, like SmartyStreets get it wrong often.
Japanese people can't have middle names (the citizen registry doesn't allow it), but foreigners can, so many systems will reject spaces in the name field. Meanwhile they're meticulous about making sure your name matches your ID exactly, leading to the situation I had.
The bank's web signup form disallowed spaces in your name, so I wrote my name FIRSTMIDDLE. Then when they processed my application they sent me an email "Your name doesn't match your ID card! Please approve the change to 'FIRST MIDDLE'."
Explanation:
echo "é" | iconv -t utf8 -f iso8859-15
We're Dutch, and the é is part of our language, and even part of the legacy character encoding standard everyone used before Unicode's widespread adoption. This is just a matter of code that works perfect as long as all characters are part of the ASCII set, but fails on the characters that don't conveniently match between UTF-8 and ISO-8859-15.I doubt these issues will go away within even, say, twenty years.
Here in CJK territory, using the wrong encoding makes the output so obviously broken [1] that mistakes are almost always caught before hitting production.
If you happen to be using Windows configured in a foreign language the first time you start Outlook, your inbox, sent mail, etc, folders are named according to that language, and will never change, and you'll have to live with non-standard names for the folders.
At least its teaches you how to configure folders manually in most email clients.
For anyone who doesn't know OSX, this translation happens on the UI level. Typing ls in a terminal gives you the real directory name.
Yes. In some languages those are actually not "markings" but denote proper letters, like in German ä,ö,ü and ß. But even if not, like in French, it can alter the meaning of words. E.g la != là. Therefore, for most Europeans and speakers of other languages that depend on more letters than ASCII provides, it is very annoying when that is not supported properly.
However, I have made the experience in a few cases that particularly Americans have a hard time understanding this. The remark about your wife not caring seems to be in this vein, too. Recently, I decided to convert our MySQL DB tables from latin1 to UTF8. (I wasn't even aware that we didn't have some form unicode, as our DB is only few years old, and I thought some unicode is the default nowadays everywhere. But then MySQL...)
Anyway, my CEO (also an American incidentally) was trying to keep me from it because he thought it's not high priority. However, we're about to go live in a French-speaking region, but which also has other indigenous languages (and therefore names), with their own "special" characters (I put "special" in quotes because for those languages, they're not "special" at all -- but I guess you get my gist by now).
Also, in previous jobs I have converted legacy systems to unicode and know what a pain it is down the road. Not to mention all the hard-to-find bugs if you don't do it, because some strings don't compare as they should, or people are just annoyed because their name is not shown correctly.
So I went ahead with the conversion anyway. We may never know for sure, but I'm convinced that I saved us some major customer frustrations, days of bug hunting and weeks of converting everything later, when existing data would need to be migrated.
So please everyone, just use UTF8 or some other unicode variant from the get-go. The few bits you might save otherwise are just not worth it.
I've been writing code to clean up a 2013 database dump. The database stored everything in LATIN-1 fields. Not because the data is in LATIN-1, but because LATIN-1 will accept any byte value. This makes error messages during input go away. See this bad advice on Stack Overflow.[1]
Some of the data is ASCII. Some is UTF-8. Some is Windows-1252. Some data is none of those, but is mostly ASCII except that there's a 0x9d once in a while. (Still haven't figured out what character set that is. From context, the ™ or ® symbol is intended.) So I have recognizers for these cases, and convert everything to UTF-8, testing every field value individually.
One column has garbaged non-English names. Someone had tried to "normalize" UTF-8 to lower case by using an ASCII lowercasing function on UTF-8 stored in a LATIN-1 field:
KACMAZLAR MEKANİK -> kacmazlar mekanä°k
Anita Calçados -> anita calã§ados
Felfria Resor för att Koh Lanta -> felfria resor fã¶r att koh lanta
I have the un-"normalized" form and can fix this.[1] https://stackoverflow.com/questions/44251813/unicodedecodeer...
There are a lot of Unicode-hostile environments out there. Java is old enough to always require explicit encoding declaration for pretty much any tool ... compiler, documentation generator, etc. Forget it at any one point and you get garbage. Reading or writing text files should always make the encoding explicit, but rarely does so. C#'s string methods all support, but don't require, a Culture parameter, without which you're practically guaranteed to do things like case conversion, or substring searches wrong in the general case. There was an awesome and long answer by tchrist on SO once about what the Perl boilerplate is to properly support Unicode for many or most circumstances (it's complicated and long and I doubt many people are going those lengths).
Point being, even when using something that supports Unicode well, the programmer still has to care, simply because text and language are messy things and it simply isn't possible to have a magic bullet that does everything right.
Much like the printing press, I'm 100% certain that the computing (and the internet specifically) is altering human written language across the world.
It is just so much easier to avoid anything outside ASCII because you can be certain ASCII will always work - even though some awful MS Access -> CSV -> SQL -> SQL -> Excel -> SQL -> COBOL ETL pipeline. No matter what version of any software is being used.
Technology has always shaped written language and we should fight to do better but at this point it seems inevitable.
(To be clear: I'm not saying this is a good or desirable state of affairs)
Eg A Ą Å Æ Ä are all different letters in most languages, not simply a pronunciation guide. I think most European languages use at least two from that list.
[1] https://quickview.cloudapps.cisco.com/quickview/bug/CSCut083...
It seems like you'd have to do something pretty stupid at the coding level to introduce a problem with "Null" by mistake. I'm sure it happens, but not more of an occasional issue.
My best guess is that there are common old databases that did not have a first class null type where it was common practice to use the string "NULL" for that purpose. And that companies that have these old systems are proactively filtering user input to prevent causing these old system to choke... It sounds like the filter is case-insensitive, though, which would be too aggressive for the case I'm thinking of. Maybe they are (mis)using a bad word filter for this, which would tend to be aggressive.
if (lastName.toLowerCase() != null);-)
It isn't a DB issue. It's more of a front-end or middle tier issue.
In SQL standards - NULL is a "marker"/TYPE. NULL and "NULL" are two separate things. One is a null type and the other is a string type.
Or more specifically, it is a "interface" problem between RDBMs and front-end since languages handle null differently. Many languages didn't have null types and null in certain languages mean different things that "lack of information".
For example, if a database column was a nullable int column and you wanted to bring it out to the java or .net space you would have issues since "int" in java and .net are value types and not reference types. So you could assigned null to the values. Where as a string/text/varchar column you could since string in java and .net are reference types and can be null.
In some languages, checking for null means you have to convert null into a string and then compare "null" == "null".
It's a legacy of lack of Nullable types in many programming languages. With the introduction of Nullable types many of these problems went away.
More importantly, who the fuck used such monstrosities?
Turned out it was his last name, Curl, that was causing the issue. The system was throwing an exception because it interpreted the customer's entry as attempting to execute the curl command.
We ended up having him put "JR" at the end of his last name to prevent it.