Truths programmers should know about case (2018)
b-list.org
b-list.org
Having a standard mapping from numbers to little graphical ideas is great. Trying to digitalize all the world's languages from a conceptual basis of "they're like English but with quirks" seems to me to be a dodge.
In other words, this article describes the kind of thing that I think should be what the Unicode people do: map out and specify datastructures and algorithms, libraries, for each language.
That isn't to say we wouldn't want accurate representations of languages (as Unicode makes an attempt to do), but there's a difference between trying to represent a language, and trying to represent the data within a language. In some cases, the former is important, in others, the latter vastly outweighs the former, and making that easily accessible, categorized, compared, etc has value.
I think it's only even conceivable at this point because of how much communication is done online, and in abridged form. We've already adopted pidgin languages for ad-hoc communication (as in IMs, SMS, and tweets), so I don't think people would find this that hard to grapple with at this point.
* You can easily represent the common symbols in a few number of bits (7bits in case of ASCII).
* It is relatively easy to do a case-insensitive comparison.
* With the exception of numerical strings, sorting and ordering is relatively straightforward.
This allowed the early software/hardware developers to create simple systems that were good enough to be useful to consumers.
The 2 main reason are:
1) you have one Sillicon Valley in USA, so all competent people were in one location. If you wanted to start Intel/Apple/Google in Denmark you'll run out of Wozniaks even before you got going because they are spread out all over the map.
2) All of USA speaks the same language, so you can easily do complex communication with companies in other states (e.g. Intel/Gateway in Texas). Getting a French person to speak good English in the 60-70'ies was a, ahem, challenge.
There is no Sillicon Valley in Europe precisly because there is no single place, and therefore Europe miss out on the "everybody important in same place" effects.
China's government pointed to one place at the map when they wanted a SV. That's why China got Shenzhen. Europe/EU still can't do that politically.
The authors conclusion was that this turned Japan into a single-use device nation (because a nintendo doesn't need you to type) and the west much more focused on PC's. I think he even went so far as to say the iPod couldn't have been invented in Japan, because it was too dependant on a PC to get music onto it (the article was pre-iphone, or around that time)
Case-insensitive comparison only became a relevant concept in informatics because it was easy on the countries that advanced on it first.
> With the exception of numerical strings, sorting and ordering is relatively straightforward.
That's another one inverting causation. If informatics developed on a place where different kinds of sorting made sense, we would be using those different kinds.
I do think your point about the small alphabet stands. But keep in mind that it applies to all languages that took influence from the ancient Greek, what is basically only excludes South-Eastern Asia.
On the other hand, I've met people who read one of those and just stopped there and concocted some absolutely bizarro ways of "handling" the issue by coming up with patterns that attempt to simply avoid everything in the list.
I really do like this approach and think that it's great to have the jumping off point and some deep exploration all in one place. Most of us who are going to get the positive effect from "falsehoods . . ." memes probably already have, and we could do with a break from those and more articles like this one.
E.G: in Python 3, it's recommend to use str.casefold() instead of str.lower() for this:
>>> "Straße".upper().lower()
'strasse'
>>> "Straße".lower()
'straße'
>>> "Straße".casefold()
'strasse'
>>>
Depending of the context, a little str.strip() and/or a str.split() + str.join() may also help. User usually don't want to consider blanks when entering data (e.g: an address), while it's important for machines (e.g: a password).Now like @lmm said, this only normalize the codepoints. Meaning is of course, impossible to be sure that way.
Opinión => OPINION
étude -> Étude français -> FRANÇAIS
http://www.rae.es/consultas/tilde-en-las-mayusculas (link in Spanish).
Personally (I'm a nobody that doesn't always agree with the Academy) I don't like omitting the accents at all in any circumstances.