That said, if you have examples where ftfy fails in any language, please submit them! We want this tool to work well, because anything that we can't fix will cause us to have egg on our faces with a customer someday...
[1]No peeking: how many format options in Excel's "Save As" dialog, excluding Excel formats, produce a document from which Unicode can be recreated accurately?
The world is a very large place, there are many codepages in use besides latin-1 and "ISO Elbonian". All central european contries use latin-2 (1250), or cyrillic codepage (1251). Since they are all single-byte codepages, they cannot be detected by try: convert() catch: try_another_codepage() and must be distinguished statistically. LTR/RTL language and asian encoding detection is even worse. https://en.wikipedia.org/wiki/Code_page https://en.wikipedia.org/wiki/Windows_code_pages
Another python Unicode conversion module which is slightly less US/English-centric: https://github.com/buriy/python-readability
Not all text arrives in the form of unmarked bytes. HTTP gives you bytes marked with an encoding. JSON straight up gives you Unicode. Once you have that Unicode, you might notice problems with it like mojibake, and ftfy is designed for fixing that.
Like you say, encoding detection has to be done statistically. That's a great goal for another project (man, I wish chardet could do it), but once statistics get in there, it would be completely impossible to get a false positive rate as low as ftfy has.
>>> ftfy.fix_text("РґРѕСЂРѕРіРµ РР·-РїРѕРґ #футбол")
'дороге Из-под #футбол'
I've even fixed a bug that occurred in Ukrainian, based on automatic testing.