String types are fine. How about your code?
eteeselink.wordpress.com
eteeselink.wordpress.com
ByteString is a series if bytes that just may happen to be ok to print as human text for debugging. The system makes it hard for you to treat it otherwise by moving the "Char8-assuming" functions to different modules and packages which must be explicitly imported and carry warnings.
You convert between them using functions in the Text.Encoding module which may fail like "decodeUtf8'" and "encodeLatin1". There's also a slew of normalizing functions.
I really encourage anyone interested in this problem to peruse the Text and Text.ICU documentation.
http://hackage.haskell.org/package/text
http://hackage.haskell.org/package/text-0.11.3.1/docs/Data-T...
http://hackage.haskell.org/package/text-icu
http://hackage.haskell.org/package/text-icu-0.6.3.7/docs/Dat...
http://hackage.haskell.org/package/text-icu-0.6.3.7/docs/Dat...
Human language and culture is too complex, fuzzy, and variable to hardwire its rules into a programming language specification. Character boundaries and character transformations are really only the beginning of it.
Consider finding the end of an English word. Is an apostrophe at the end an actual apostrophe denoting the possessive case (of a plural word) or meant as a closing single quote (even where there are different symbols, they are often mixed up when doing data entry)? How do you pluralize words? You have to consider exceptions to the usual rules (e.g., a dictionary), words that only exist in the plural, languages that don't have the concept of plural words, etc.
If I have a database of employee bios and I write an app so HR can search that database, it's no problem for them to type in "management" and get a list of all employees with the word "management" in their bio. But when a division does the same thing in a non-English-speaking country where there's an é in the word, then that word can be composed of either two characters or one and searching one way will tend to miss any results which were composed of text the other way. The solution here is to normalize the data before it is stored in the DB and the search term before it is searched for.
It's a simple concept. Not all programmers are aware of it, however, and this guy totally missed the point.
I bet the MS Word team that made internationalized search/replace spent a lot of time getting it even somewhat right.
Everybody knows about Bobby Tables? How about Mr. เจ้าพระยาบดินทรเดชา Unicode?
What if somebody told you that if certain exotic names are entered into the system by certain input method, government officials can't find those people in their searches, including full text search to a police report?
Not actually
Just off the top of my head. There are far more. You can't even have a native language user friendly login system without normalizing text.
search, that counts
text replacement,
why?
ellipsizing text at the end or middle (as done commonly on iOS),
you need to know width for that, and you don't, bad idea
fuzzy search,
search
password word limits, password length limits, matching passwords (try using bcrypt on international languages without first normalising),
i18n password? very bad idea
detecting and filtering rude words in public chat
search&replace, anyway you won't be able to, because f͌̊ͦuͥcͥ͑ͩ͛k͂͌̿ͣ̾
You also need to not chop off accents. "café" cannot become "cafe…"
> i18n password? very bad idea
Because string types are so broken.
While I think unicode support in languages could be better; there is a lot of truth in this article that surrounds his subject core.
HTTP is a good example of something normally using string typing; most web servers treat headers, for example, as text, rather than as data, and so you end up with a lot of invalid stuff because of misunderstandings about what the values actually are or how they should be interpreted or written.
That's a trap that I'm dodging in rust-http; in there, (known) headers are data, serialised to text only for transmission, not text. (I'll freely admit, however, that this is an approach I would never take in Python; I'm confident it would lead to more errors than the alternative.)
Obviously it is likely an effect of a global design flaw, but such things are very hard to fix.
99.9% of the time I have seen code do crazy things "because it is faster" it's not performance critical code anyway and there is no explanation provided.
Additionally the XML and question was all precisely to the byte level the same until you hit the giant (100s of Megs) based 64 blob that was the content. The parser stripped X number of bytes from the start of the file, and from the end, and de-base-64ed the center - which if I recall it then sent off to another parser as the content was in some old but standard record format from the 80s.
Anyhow - I'd say using XML in this case was the abuse, not the substring. But we were in no position to get the vendor to change their data format so...
You can slap custom elements pretty much anywhere you want, as long as you have your own namespace (and it's recommended you only place them under <message> or <iq> elements). Say you have some proprietary technology in a client application, with XMPP you can throw an element under the <message> that your client can recognise and act on. For everyone else provide a hyperlink within the <body> element and serve up a web page for them. If they are using your client "bam!" instant added functionality - but if they are on device X which you do not support they are not left out in the cold.
E.g. you can sometimes skip over chunks of characters without every accessing them, and get speedups of magnitudes over even "just" checking every byte in the input.
There's nothing a faster proper XML parser can do about a custom parser like that.
(Obviously this is a brittle solution and a last resort optimisation, and should be accompanied by ample warnings, but sometimes there are no alternatives)
string DoStuff(string arg)
Where the service call takes a string that actually contains XML and returns a string that actually contains XML....My own "Worst use of XML" was someone wanting to use quite complex XML documents as a key in a database....
The original article was wrong because it proposed replacing strings with arrays of code points. Clearly, that doesn't work.
This article is wrong because adding more string types just shuffles the problem around. There is nothing "machine consumable" about strings encoded in a certain anglo-centric character encoding! Just don't even think that thought. "abc" is absolutely not more "machine consumable" than "東京". You don't "hash prose" by transliterating text into ascii characters.
It's not impossible to fix the existing string types. Principle of least surprise holds. In cases when it doesn't, the locale is the tie-breaker. E.g. "Scheiße".upper(locale.DE) may be different from "Scheiße".upper(locale.RU).
"Łódź" may be the same as "Lodz" to you in the same way that "komputor" may be the same as "computer" to a Russian speaker. To someone else the names "Anderson" and "Andersson" are equivalent. Now you see the problem -- exact matching is futile and you should use fuzzy matching instead, like normalized Levenshtein distance, and rank the results based on similarity.
Even that is not enough if you want to support non-Latin alphabets because they have different ideas about what a character is but it should get you started.
In some theoretical sense, yes. In terms of solving users' problems and providing business value, I need to make it possible for users to find the entry for "Łódź" without typing accents.
> (What language are we talking about, anyway?)
We're talking about Scala; it runs on the JVM and so java.lang.String is the string type.
This thread is talking about two completely different types of Equals; we shouldn't be using the same word for them and certainly not the same function name in code:
- SomeString1.Equals(SomeString2) -> are all the bytes in array 1 equal to the bytes in array 2?
- HumanSimilarity(t1, t2) - given text s1 and text s2, give me a number that tells me how similar a person would perceive these strings to be. You could even go further:
SimilarityForReaderInLocale(locale, t1, locale1, t2, locale2) - for a human reader in locale, given text t1 written by a human in locale1 and text t2 written by a human in locale2, how similar would the reader perceive these two pieces of text?
When we talk about 'least surprise' in manipulations, what I think we really mean is that text manipulation should be defined by a locale; no actually, a superset of a locale. That superset being your typical human reader in that locale.
Human language? English speakers talking about cities in a variety of countries (which should be correctly named, but searchable by english speakers).
Sure, but the language should offer support for this, even if it requires some level of configuration. It's not a problem we want every programmer solving anew for every program.
> I wouldn't rely on something like string.equals for it.
True enough. But I think the existing String.equals method is broken: it behaves very surprisingly and leads programmers to introduce bugs. Likewise e.g. String.subString (which can chop a character in half). These methods cannot be fixed (because existing code assumes their current functionality) and are very difficult to use effectively; they should be deprecated, which in practice means a new String type.
Yeah, I completely agree. You can make an argument that this stuff belongs in a library (or even several libraries) rather than the language, simply because of how many decisions are involved:
- do you know what language your search input is in? can it be any of several languages? any at all? only one specific language, ever?
- what language is the text you're searching on? do you know? are you potentially searching across text in multiple languages at once?
- how exact does the match need to be? do you want phonetic matches? do you want to match characters that "look" similar but many sound different?
- do you need a binary or a fuzzy match? (e.g. are you doing a search ranked by relevance). Do you need to compute some sort of Levenshtein distance?
So the use cases can range from a simple byte-level exact comparison all the way to a full search server like Solr. How much of that should the language implement? (not a rhetorical question, I actually don't think there's an obvious answer).
No arguments on the existing String methods :)
> If there is any takeaway from this entire discussion, it may be that there is a need for multiple string types in strongly-typed languages
Yes, yes there is. Until we get that, string types are not fine.
It is my opinion, however, that string types are fine, just not perfect. I should have maybe made that clearer.
The whole problem is that current string types enable broken unsafe behavior on Unicode ("human only" in your parlance) strings. Current string types are broken because they do not enforce the requirement that string operations are done only on plain ASCII ("machine only") strings.
Calling ASCII "machine only" is totally wrong. I agree that encoding/decoding is a pain, but we have different string encodings for a reason.
In my experience, you need to do this any time you're displaying someone's name, a place name, an article title, or whatever. Often the display area just isn't that wide, and shifting around other content may not be an option. You need to display enough to let people know what's there, but eventually it needs to be cut off.
In particular, problems arise when some function renders a sequence of characters contained in a string [fundamental string operations such as concatenation tend not to be problematic]. These problems are due to the transition from the mathematical certainty of strings to the heuristics of text rendering. The compromises required to map semantics of human writing systems onto strings via Unicode contributes to this problem...glyphs are not necessarily ordered sequentially or without resolution under a context.
Nope.
If anything, change 'probably' to 'might'. Confidence is good, but some people could take 'probably' as an imputation on their ability to comprehend the subject matter. I'd always aim at "That was just the beginning. Whip out your chequebook and then you'll REALLY see what I have to offer" rather than "You sound like you're in trouble."
Thanks for the support, but if I make an arrogant remark, I have to expect to get snarky responses, right :-)
The line was supposed to be taken as a joke, in reference to people signing off their blog posts with "If you read this far, you should probably follow me on Twitter" and the likes.
Thanks for your feedback though! I hadn't thought about how people could interpret it as a slight insult to their understanding of the subject matter, so next time I make an arrogant joke I'll try and take stuff like that into account.