Python-ftfy: Given Unicode text, make its representation possibly less broken
github.com
github.com
>>> ftfy.fix_text('López')
'López'
Bravo!Note: this module, like UnicodeDammit, is very US/English-centric, and is practically useless for worldwide web scraping. For non-english pages, it is necessary to statistically estimate the codepage and language of each page segment, and then try to normalize each segment to unicode.
The world is a very large place, there are many codepages in use besides latin-1 and "ISO Elbonian". All central european contries use latin-2 (1250), or cyrillic codepage (1251). Since they are all single-byte codepages, they cannot be detected by try: convert() catch: try_another_codepage() and must be distinguished statistically. LTR/RTL language and asian encoding detection is even worse. https://en.wikipedia.org/wiki/Code_page https://en.wikipedia.org/wiki/Windows_code_pages
Another python Unicode conversion module which is slightly less US/English-centric: https://github.com/buriy/python-readability
Not all text arrives in the form of unmarked bytes. HTTP gives you bytes marked with an encoding. JSON straight up gives you Unicode. Once you have that Unicode, you might notice problems with it like mojibake, and ftfy is designed for fixing that.
Like you say, encoding detection has to be done statistically. That's a great goal for another project (man, I wish chardet could do it), but once statistics get in there, it would be completely impossible to get a false positive rate as low as ftfy has.
That said, if you have examples where ftfy fails in any language, please submit them! We want this tool to work well, because anything that we can't fix will cause us to have egg on our faces with a customer someday...
[1]No peeking: how many format options in Excel's "Save As" dialog, excluding Excel formats, produce a document from which Unicode can be recreated accurately?
>>> ftfy.fix_text("РґРѕСЂРѕРіРµ РР·-РїРѕРґ #футбол")
'дороге Из-под #футбол'
I've even fixed a bug that occurred in Ukrainian, based on automatic testing.You should perhaps take a look at the "sloppy-windows-1252" codec in ftfy, and it may help "detwingle" handle some messier cases. (For example, Python will say 0x81 isn't a valid byte in Windows-1252. It's technically right. But there it is anyway.)
Empirically, it's not in ISO-8859-2.
I think the problem here is that chardet is built on the assumption that "encoding detection is language detection" (from its docs). This assumption is necessary, and basically correct, when distinguishing Japanese encodings from Chinese encodings. It's even pretty much taken as a given that you can't have Japanese and Chinese text in the same document without contortions that most developers are unwilling to go through.
But European languages and encodings are much more intermixed than that. One document may contain multiple European languages, and these languages may be written outside of their traditional encoding.
I wouldn't know how to fix the European languages without damaging chardet's clear success at distinguishing East Asian encodings.
http://www.whatwg.org/C#determining-the-character-encoding is what the spec defines, and though the eventual fallback is implementation-defined (Firefox, for example, combines locale-specific defaults with TLD-based defaults).
But I built a sanitizer in a couple hours with this lib, and it seems to work pretty well.
The only unexpected thing is that it converts the ordinal indicator º to o in addresses. Luckily there are only a handful I need to fix.
>>> print ftfy.fix_text(u'ordinal indicator º to o in addresses.')
ordinal indicator o to o in addresses.
>>> print ftfy.fix_text(u'ordinal indicator º to o in addresses.',normalization='NFC')
ordinal indicator º to o in addresses.Let me explain in some detail why this library is not a good thing:
In an ideal world, you would know what encoding bytes are in and could therefore decode them explicitly using the known correct encoding, and this library would be redundant.
If instead, as is often the case in the real world, the coding is unknown, there exists the question of how to resolve the numerous ambiguities which result. A library such as this would have to guess what encoding to use in each specific instance, and the choices it ideally should make are extremely dependent on the circumstances and even the immediate context. As it is, the library is hard-coded with some specific algorithms to choose some encodings over others, and if those assumptions does not match your use case exactly, the library will corrupt your data.
A much better solution would perhaps involve a machine learning solution to the problem, and having the library be trained to deduce the probable encodings from a large set of example data from each user’s individual use case. Even these will occasionally be wrong, but at least it would be the best we could do with unknown encodings without resorting to manual processing.
However, a one-size-fits-all “solution” such as this is merely giving people a further excuse to keep not caring about encodings, to pretend that encodings can be “detected”, and that there exists such a thing as “plain text”.
It's the library you use when the data you get has already been decoded incorrectly. The user of ftfy cares about encodings, but gets data from sources that don't.
And in no practical sense does it corrupt your data. I don't know where you got that idea from. It leaves good data alone.
I will not say that false positives are nonexistent, but they are vanishingly rare -- see http://ftfy.readthedocs.org/en/latest/#accuracy -- and they don't occur in "serious" data, they occur when people are screwing around with bizarre emoticons and stuff.
Your protestations about how corruptions will not occur in “serious” data (what is is that, anyway?), and blaming “bizarre emoticons” is exactly symptomatic of what I’m talking about – you are blaming every wrong guess which this library makes on uses of non-ASCII. This is being an ASCII neanderthal. The problems of ignoring encodings are real, and should not be blamed on users of “bizarre emoticons and stuff”.
You could vaguely criticize the fix_entities='auto' setting as a guess, except it's a guess that's only wrong if you manage to provide it an HTML document with zero tags in it.
An example of a false positive is "├┤a┼┐a┼┐a┼┐a┼┐a". That is what I mean by non-serious text. False positives will always exist, and you should appreciate that I'm testing on millions of examples to find out what they are. Your suggested machine learning approach would never get to 99.999984% precision.
Using non-ASCII is totally fine, and this library would have no purpose in an all-ASCII world.
Another commenter claims that this worked:
> ftfy.fix_text('López')
There is no “fix_entities = True” there. Indeed, if the library would require such a parameter, what would be the point? If you already know the encoding, the library has no reason to exist. Therefore, the whole point of the library is to guess the encoding.
> you should appreciate that I'm testing on millions of examples
Examples taken, I would assume, either from your personal use cases, the use cases of your customers, or some sort of general grab-bag of mis-encoded text you could find. I would assume that this one-size-fits-all ad-hoc rule set would be wrong for many users in their specialized use cases, and would bite them when they least expect it.
You're not even reading the documentation, you're just searching for reasons to call me an "ASCII neanderthal" over a library that an ASCII neanderthal would have no use for.
And I fail to see how the default settings being able to fix 'López' is anything but a resounding success.
Anyway, now we’re just arguing in circles. I said that the library would guess that “&” was HTML encoded. You said “fix_entities is a parameter. It's not guessing, you told it.” I said that the example given has no such parameter. You then turned around and said that it was a success that it guessed correctly, but my point was that it was, indeed, guessing, and might, therefore, guess wrong.
I don’t want to call you an ASCII neanderthal, really, and I’m sorry I did, it’s just that your library helps the actual ASCII neanderthals from having to bother with evolving. This is my main complaint about this library. It will probably be used by them, reflexively, to decode everything, even when the encoding is known, and therefore introduce (admittedly relatively small amounts of) data corruption (but these things have a tendency to crop up when you least wanted them). Whereas if your library was not used, users would have to think about what decoding to use, use it, and not introduce data corruption.
Also, I have some misgivings about a one-size-fits-all solution of guessing encodings – I guess that it would never really work in quite the painless way most your users imagine. To solve this, I advocated a user-customizable training approach, which would, for each user, be the best possible one for their use case. It would also have the beneficial side effect of forcing the users to actually think about their data and what encodings it was likely to have and in what circumstances, thus making them evolve to Homo Unicodus. ☺ Of course, I could be wrong about this, but my principal worry about this library, as stated above, remains.
Keep in mind that this is a library that finds emoji that has been damaged by "serious" software written by, let's call them "Basic Multilingual Plane neanderthals", and puts it back. There's your fun.
This library is taking the only possible approach, which is to segment the text and try to convert each segment into its most probable unicode representation. It seems to cover a larger number of encoding mixups compared to other libraries, that's great!
Comment to rspeer: in fix_text_segment(), I would limit the number of recursive passes on the text to 5-10. Right now it's using 'while True', which might take a very long time to converge on corrupted/binary data.
I have also seen text that was encoded six times in UTF-8 (and decoded five times in Windows-1252). Although ftfy had to leave it as is; it didn't successfully decode because it was truncated.
It is taking the only possible approach if we assume that it must use one and only one algorithm for all uses. Otherwise, it seems to me that a lot of careful tuning and configuration would be needed in order for this library to make the best guesses it possibly can make for a specific user’s situation and data.
> might take a very long time to converge
A limit there might be appropriate – otherwise there might exist a “billion laughs” style attack.
Wait, does this not make the usability of this library very limited? How would I know that something has been decoded incorrectly? If I’m already handling this manually, what is the point of having a library?
Your questions could be answered, but not by me. Plonk.
Their approach is then to attempt each type of encoding they got, ignoring errors, and show the result to the user so the user can decide (detect) which decoding worked and which didn't.
Can ftfy replace this functionality? Is it doing the very encoding detection which is currently done by humans?
There's some auto-detection you can do -- for example, you can distinguish UTF-8 from byte-order-marked UTF-16 with 100% accuracy, by design. You could also try chardet if you're okay with some amount of errors. Maybe show the detected encoding first.
Cases where ftfy is useful:
* Web scraping -- sometimes you get data that decodes in the encoding it claims to be in, but isn't quite right
* Handling data that has, at one point, been imported and exported in Microsoft Office, without every user consistently picking exactly the right format from like 20 inaccurately-named options
* Handling data that was stored by half-assed services written in, say, PHP, and not tested outside of ASCII
* Reading CESU-8 (the non-standard encoding that Java and MySQL call "UTF-8", for backward compatibility) in Python without breaking it even more. (This isn't automatic.)
* Handling data that's combined from multiple sources, in mixed encodings
* All the other situations in which mojibake arises in mostly-readable text, which there seem to be no end of.
1. Due to its simplicity for a large group of naïve users, the library will likely be prone to over- and misuse. Since the library uses guessing as its method of decoding, and by definition a guess may be wrong, this will lead to some unnecessary data corruption in situations where use of this library (and the resulting data corruption) was not actually needed.
2. The library uses a one-size-fits-all model in the area of guessing encodings and language. This has historically proven to be less than a good idea, since different users in different situations use different data and encodings, and your library’s algorithm will not fit all situations equally well. I suggested that a more tunable and customizable approach would indeed be the best one could do in the cases where the encoding is actually not known. (This minor complexity in use of the library would also have the benefit of discouraging overuse in unwarranted situations, thus also resolving the first point, above.)
You have, as far as I can see, not yet responded substantively to the first point, and for the second point you have only asserted that your user-uncustomizable alorithm is superior to any possible other automatically derived algorithm.
Yet, I’m the one who deserves a Plonk? I think not.
Which leads me to my other concern: Why do you use NFKC compatibility as the default normalization? Given you are a text mining company, you of all guys should know you loose valuable information - particularly about numbers, super- and subscript characters - with this normalization strategy. Doing NFKC on stuff like all kinds of articles, books, patents, etc. would lead to potentially disastrous results (e.g., NFKC "decomposes" the string 'O\u2082\u00B9' to 'O21' instead of 'O_2^1' - "oxygen, reference 1"). In general, I think NFC is what Python and many other libraries do, while I believe NFKC should only be used when you know what you are doing (and why you need it). Maybe it is useful for some strange, geeky tweets, but I would argue that its the corner case, not the default.
When it comes to text analytics, the underlying tagger and stuff won't know what O21 is any more than it knows what O_2^1 is anyway. And NFKC is useful for mixed Latin and Japanese text, which I wouldn't entirely dismiss as strange and geeky. But it's true that the default could be more conservative.