UTF-8 Everywhere (2012)
utf8everywhere.org
utf8everywhere.org
You rarely need to index a string with an integer in Python. FOR loops don't need to. Regular expressions don't need to. Operations that return a position into the string could return an opaque type which acts as a string index. That type should support adding and subtracting integers (at least +1 and -1) by progressing through the string. That would take care of most of the use cases. Attempts to index a string with an int would generate index arrays internally. (Or, for short strings, just start at the beginning every time and count.)
Windows and Java have big problems. They really are 16-bit char based. It's not Java's fault; they standardized when Unicode was 16 bits.
I agree, however, that it's completely irrelevant whether the indexes correspond to code units (i.e. byte offsets in UTF-8) or whether they correspond to code points (how it works in Python currently), as long as we have some way to store, compare, and otherwise manipulate locations within a string.
Some Rust developers at one point proposed making string indexes their own (opaque) type, as you suggest, so that they couldn't be confused with integers used for other purposes. The extra complexity of such an API meant that this proposal was never really taken seriously, and it only prevents a small category of programming errors.
You might be interested in looking at some string APIs which are mostly without string indexing, like Haskell's Data.Text, which is one of the most well-designed string APIs ever made.
https://hackage.haskell.org/package/text-1.2.2.1/docs/Data-T...
As for Windows, my Windows apps use UTF-8 everywhere, and then convert to wchar_t at the last possible moment when interacting with the Windows API. I believe this is what UTF-8 Everywhere suggests.
That's why I really like dealing with UTF-8. As you said you can just index by byte instead of having to worry about code point boundries. This is because a code point will never match inside a larger code point.
So if I search for a one byte quote character it will never match the second, third or fourth bytes in a larger code point. Same with other control characters.
Works when searching for 2 or 3 byte code points as well.
If you want to split the string into "characters" you need to do it a the grapheme level (multiple code points) and for that you need to use a unicode library. But that adds overhead when you're just scanning.
A JSON or XML scanner does not need the added overhead of advancing by code point or grapheme.
For parsing, it often makes sense to iterate over code points or code units, since many languages are defined in terms of code points (and you can translate that to code units, for performance). XML 1.1, JavaScript, Haskell, etc... many languages are defined in terms of the underlying code points and their character classes in the Unicode standard. JSON and XML 1.0 are not everything.
For parsing it's easier to just scan for a byte sequence in UTF-8 because you know what you're looking for ahead of time. If you're looking for a matching quote, brace, etc. you just need to scan for a single byte in your text stream. Adding a smart iterator to the process that moves to the start of each code unit is not necessary and will slow things way down.
I just gave JSON and XML as examples and not an exhaustive list. If you know the code points you are scanning for it's way more effecient to scan for their code units. The state machine in a paraer will be operating at the byte level anyways.
I have yet to see a good example where processing/iterating by code point is the better choice (other than the grapheme code of the unicode library).
If you're curious, here's the V8 tokenizer header file:
https://github.com/v8/v8/blob/master/src/parsing/scanner.h
You can see that it works on an underlying UTF-16 code unit stream which is then composed into code points before tokenization. This extra step with UTF-16 is a quirk of JavaScript.
If you think that V8 shouldn't be processing by code point, feel free to explain that to them.
For languages that only allow non ascii in string literals a pure state machine would suffice.
Not sure why you're mentioning parsers. At that point you you're dealing with tokens.
As for UTF-16 it's an ugly hack that never should have existed in the first place. Unfortunately the unicode people had to fix their UCS-2 mistake.
Since Javascript is standardised to be either UCS-2 or UTF-16 it probably made sense to make the scanner use UTF-16.
ECMAScript source text is represented as a sequence of characters in the Unicode character encoding, version 3.0 or later. [...] ECMAScript source text is assumed to be a sequence of 16-bit code units for the purposes of this specification. [...] If an actual source text is encoded in a form other than 16-bit code units it must be processed as if it was first converted to UTF-16.
Really I think you are arguing against the notion of "default iteration" altogether. As you say, the right type of iteration is context dependent, and it ought to be made explicit.
In fact it looks like even standard indexes have been deprecated, in favor of Iterators over the string:
https://doc.rust-lang.org/stable/std/primitive.str.html#meth...
Iteration is better IMO.
I've been using references pretty effectively to get around some issues like this, though in other cases Rc is my only resort.
If you have some code up on github, I'd be happy to take a look and see if there's some other options that are less cumbersome.
I think Python got unicode in the same era as Java, so it's understandable that Python 2 doesn't work like this. But if they are going to break the whole world for unicode, I also think it would have been better to do something like Go does (e.g. the rune library).
Edit: see pjscott's comment below - it's by code point, not byte, but still not by character.
$ python3
Python 3.4.3 (default, Oct 14 2015, 20:28:29)
>>> s = "オンライン"
>>> s
'オンライン'
>>> len(s)
5
>>> s [0]
'オ'
>>> s [1]
'ン'
>>> s [2]
'ラ'
>>> s[3]
'イ'
>>> s[4]
'ン' $ python3
>>> "위키백과"[1]
'키'
>>> "위키백과"[1] # Should be identical, right?
'ᅱ'That is perhaps the most succinct and accurate way I've heard to explain and justify why you're sounding like a wet blanket to people that may not understand, while acknowledging that you know how you sound, but there is a reason for it. I expect to use this in the future.
"difficult"[2]?
Is it 'f'? Or the ligature `ffi`? :)c = "é"
c[0], c[1]
It's the same phenomenon with Latin characters. (Extra bonus: for me, the combining acute accent character then combines in the terminal with the apostrophe that Python uses to delimit the string!)
Another idea to see the effect is "a" + "é"[1]. (The result is 'á'... and as in your examples, a precomposed "̈́é" is also available which doesn't exhibit any of these phenomena.)
This doesn't matter much for (normalized) western European text, but if the language in question needs to use separate diacritical code points you'll likely end up with hanging accents in the like. Swift is the only language I know of that has grapheme clusters as the default unit of character, I'd love to see it in more places.
[1]: http://unicode.org/reports/tr29/#Grapheme_Cluster_Boundaries
But, Perl 6 definitely gets it more right than other implementations I've seen.
Moving to grapheme cluster boundaries means that the algorithm may work incorrectly if you input a string of Unicode N+1 to an implementation that only supports Unicode N. It also makes the "increment character" function very complicated. In the UTF-8 version, this looks roughly like:
char *advance(char *str) {
uint8_t c = (uint8_t)*str;
/* Count the number of leading 1's */
int num1s = __builtin_clz(~c) - 24;
if (num1s == 0) return str + 1;
return str + num1s;
}
Grapheme-based indexing looks like this: char *advance_grapheme(char *str) {
while (true) {
uint32_t codepoint = read_codepoint(str);
str = advance(str);
uint32_t nextCodepoint = read_codepoint(str);
/* This is typically something like table[table2[codepoint >> 4] * 16 + codepoint & 15]; */
GraphemeClusterBreak left = lookupProp(codepoint);
GraphemeClusterBreak right = lookupProp(nextCodepoint);
/* Several rules based on left versus right... */
}
return str;
}
See the vast difference in the two implementations? It's a lot of complexity, and it's worth asking if that complexity needs to be built into the main library (strings are a fundamental datatype in any language). It's also important to note that it's questionable whether such a feature implemented by default is going to actually fix naive programmers' code--if you read UTR #29 carefully, you'll notice that something like क्ष will consist of two grapheme clusters (क् and ष), which is arguably incorrect. Internationalization is often tied heavily to GUI and, especially for problems like grapheme clusters, it arguably makes more sense for toolkits to implement and deal with the problems themselves and provide things like "text input widget" primitives to programmers rather than encouraging users to try to implement it themselves.I with mtviewdave that more complexity means this should really be in the std lib. Java's charAt(int i) is misleading at best.
(Probably a bit "unfair" to "pounce" on an off-hand paranthetical like this, but I'm in a bit of a pedantic mood...)
This is not true for e.g. Haskell. In Haskell it's defined as [Char], i.e. a list of characters. (Of course the Haskell community is suffering from that decision, but that's another story.)
I'm not sure why strings would need to be a fundamendal type, though. Sure, they would probably be part of the standard library for almost all languages, but they don't need to be "magical" in the way most fundamental types (int, etc.) are.
However... How does a language or library attempting to abstract this part (like swift might) deal with the other, unrelated annoying aspects of Unicode? Even if 1 glyph is 1 "char", and we normalize all the inputs, there is still, say ... Bi-di text.
PS: bit of trivia that people forget these days, Win32 has been supporting it for longer than android/iOS.
What i find more frustrating is how the documentation for many systems describes the basic unit of text as a character, without specifying whether a code point or grapheme is meant, and without leading people to an explanation of the difference. There is still a lot of software that processes unicode text incorrectly, not because it is difficult to do so, but because nobody told the developer how things should be done.
My point is that I don't think throwing people into the deep end and expecting them to grok the codepoint/grapheme division before they get to use the language is likely to be productive. Defaulting to graphemes carries the advantage that if a programmer who doesn't know much about languages that require more Unicode finesse uses it purely intuitively, they'll get things right in a lot of cases. Using codepoints, on the other hand, makes it easy to put text through a grinder while doing relatively innocent things.
Unfortunately while we're discussing the finer points of using graphemes and codepoints in APIs a million supplementary characters were brutally eviscerated by code running on languages that haven't quite gotten past UCS-2 XD
Really, this is not dissimilar to "what do you mean, there are more letters than A to Z?" issue that plagued software written in US back before Unicode became dominant. The way we (my perspective on this is as a native speaker of a language with a non-Latin alphabet) have eventually solved it is by basically forcing Unicode onto those people. It broke their simple and convenient picture of the world, and replaced it with something much more complicated. But it was necessary.
My position is that letting programmers get away with a simplistic view of text processing (by allowing defaults that "mostly" work) is what creates those issues. So adjusting the abstractions such that they expose more of the underlying complexity is a good thing. People SHOULD believe that doing text processing the right way is hard, because it is.
Also Elixir and Perl 6. For a bit more info see https://news.ycombinator.com/item?id=
It's especially amusing because on Python 3 strings internally cache their utf-8 equivalent if it was used once.
https://github.com/python/cpython/blob/master/Include/unicod...
Or throw an invalid index exception if the top to bits are one if that makes more sense for the language you're using.
So what you're saying, is that it's very easy to get corrupted strings by anyone who doesn't have an understanding of utf-8 at the bit-level - which in my experience seems to be the majority of programmers.
And indexing by code points doesn't solve the problem either. The majority of programmers don't know what a grapheme is or how to collate or sort unicode strings.
My only gripe with your argument is that I don't think it's easy to avoid corrupted text in modern text processing - which is precisely why there are libraries for it because it's actually really easy to get it wrong - even if you know what you're doing.
When I was working on a project before Unicode we would switch our dev PCs to the other languages we supported. What a pain that was. Only issues we had was when a translated string was much longer than the screen space allocated to it. I belive Swedish was the main culprit. No problems with simplified and traditional Chinese as those were more compact. I have no sympathy for dev shops that can't get internationalization right. As with everything else in the corporate dev world management doesn't seem to want to hire/retain the more experienced programmers.
I think you have a gripe with my argument because you may be missing my point. If a high level language chooses to let a programmer index into a UTF-8 string at the byte level (for performance and other reasons) it's very easy for it to prevent the the programmer from slicing in the middle of a code unit.
The reason being is that the language function to slice a unicode string would either throw an exception or just advance to the next valid index. There wouldn't be a way for the programmer to slice a unicode string in the middle of a code unit.
I get your point, it just doesn't apply to many real world situations I've seen where you don't have the luxury of just using a higher level language or a library that takes care of all these things, or keeping programmers who don't understand what they are doing away from that sort of thing.
The most egregious example that I've personally seen was a developer working on a legacy Cobol banking program that needed Chinese support retro-fitted to it.
The app was originally only developed with ASCII in mind and so sliced through strings willy-nilly, which naturally caused problems with Chinese text.
The developer working on the "fix" before me, was calling out to ICU through the C API of the version of Cobol that we used and was still messing things up - he'd actually modified ICU in some custom way to prevent the bug from crashing the program, but was still causing corrupted text.
I basically undid all his changes, and wrapped all COBOL string splicing to call a function that always split a string at a valid position - truncating invalid bytes at the start/end as necessary. Much simpler and resulted in the removal of an unnecessary dependency on ICU.
This bug had been outstanding for several months when I first joined that company, and it was the first one I was assigned to work on - and luckily for them they'd accidentally hired someone who had done lots of multilingual programming before.
it's very easy for it to prevent the the programmer from slicing in the middle of a code unit.
Okay, but even you made a mistake in your first example of what to do, and that's the sort of code that someone who knows what they are doing could write, and will seem to work in the conditions under which it was tested (working on my machine, ship it!), but that will cause seemingly random problems once it hits users.
No, I still think your missing some of it. I am not advocating that what I said is the solution for everything.
Someone said that slicing UTF-8 strings leads to string corruption and endorsed the Python 3 frankenstien unicode type as a way to avoid it. I just gave a way of preventing that.
Now you argued that a novice programmer would fail to implement it properly. So you're comparing my method implemented by a novice programmer to a method implemented by profesional compiler writers. That hardly seems fair. :)
So my argument is that if my method were to be implemented by professional compiler writers it would prevent corrupted strings while still using UTF-8 as the internal representation.
> I basically undid all his changes, and wrapped all COBOL string splicing to call a function that always split a string at a valid position - truncating invalid bytes at the start/end as necessary.
> luckily for them they'd accidentally hired someone who had done lots of multilingual programming before.
So an expert programmer implemented a string splitting function that didn't corrupt strings. :D
> but even you made a mistake in your first example of what to do
I writing this on an iPad while watching TV and playing a game on another android tablet while looking at the wikipedia UTF-8 article on a tiny phone screen while a little white dog is trying to bite my fingers (wish I was making this up). Not exactly my usual programming environment. ;)
sigh if only it was novice programmers making these mistakes :-/
The stuff I've seen in some people's multithreaded code just makes me want to cry.
A "string" means "a sequence of characters". Wide characters (or the equivalent interface) preserve this property.
Graphemes operate at a higher level than characters. You could construct a grapheme-strings, I suppose, but that has tons of edge cases, and if you don't like character-strings, I doubt you will like grapheme-strings.
If we chop UTF-8, we can end up with bad characters, or possibly invalid overlong forms.
If you chop in the middel of a code unit then you end up with U-FFFDs. In both cases the visual representation has been altered.
As I wrote elsewhere it is easy for the slice routines of a language to check to see if the programmer tried to slice in the middle of a code unit and either return an error or just advance to the start of a code unit.
Slicing a code-point-character string destroys only graphemes.
Clear win.
A code point string has other niceties, like being indexed by simple integers. If end is the index of the last code point of a grapheme, then the next grapheme starts at end + 1.
If end is the index of the last UTF-8 encoding of a code point, then the next grapheme does not start at end + 1.
We can have it so that it does by making end point to the last byte of the UTF-8 encoding of the code point; but then it doesn't point at the start of the character, recovering which is awkward.
The code uglification can be addressed by piling on abstractions: integer-like iteration gizmos that can be incremented and decremented thanks to function or operator overloading.
I feel that that level of abstraction has no place in character-level data processing, if anywhere, whose basic operations should be expressible tersely in a few machine instructions.
Also, we mustn't lose sight of what the T means in UTF-8: transfer. It's not called UPF-8 (the Unicode processing format in 8 bits).
Working with UTF-8 instead of with the objects that UTF-8 denotes is like working with a textual representation of Lisp s-expressions that still contain the parentheses and whitespace delimitation, and quotes around strings and so on, refusing to parse them to obtain the object which they represent. People who do this should immediately turn in their CS degrees.
All those other issues you refer to are addressed by more parsing. If you want the glyphs, the correct thing is to parse the code-point string and make a list or vector of glyph representations.
With that representation you can still break the text "carpet" into "car" "pet" which destroys semantics; that is dealt with by parsing into words.
Chopping lists of words destroys phrases; so parse phrases, and transform at the phrase level.
And so on.
I'd call it a slight improvement. And after the major step back of using 2 to 4 times more memory for strings I'd call it a net loss.
> A code point string has other niceties, like being indexed by simple integers.
Again, no benefit of this. The only argument I've heard here is to prevent bad slicing and I've shown a way to prevent that.
> Also, we mustn't lose sight of what the T means in UTF-8: transfer. It's not called UPF-8 (the Unicode processing format in 8 bits).
By this argument we can't use UTF-16 or UTF-32 for internal processing of strings either. Back to code pages then.
Many unicode languages work with UTF-8 or UTF-16 internally. So working with the "transfer format" is common practice.
While it may not be necessary to know who the languages your program in work under the hood, expert programmers do want/need to know. That way they can write better code, or switch to another language or get the language devs to improve their internal handling.
The programmer shouldn't have to know that a newline character is written as \n in a JSON string.
The JSON string "a\nb" take 6 characters to write, but it's length should be given as 3.
99% people want to manipulate a JSON model, not the JSON (or BSON) serialization itself. The 1% can still use a byte array and do whatever hacks they like.
A better example is if you want to find a newline in a string. If you do a find it in a UTF-16 string it may be position 8 and a find in UTF-8 may be position 12. Does it matter what the actual number is? NO. You just pass it to the next function or whatever.
The term character has many meanings. Graphemes are characters and that's what most users expect, something that's displayed as a single graphical unit.
"As with glyphs, there is no one-to-one relationship between characters and code points. What an end-user thinks of as a single character (grapheme) may in fact be represented by multiple code points; conversely, a single code point may correspond to multiple characters."
All you need to do is index/slice a string half-way through any character that is outside Unicode's Basic Multilingual Plane
I get that this might seem pedantic, but it's important to be pedantic about this, otherwise misconceptions and ambiguities occur e.g. 'just use wide characters' - the definition of which changes depending on the platform.
Second of all, "wide enough" for all intents and purposes means 32 bits. Technically Unicode only needs 21 bits to cover the currently defined codespace, but computers don't deal well with that and so 32bits is the minimum "wide enough" character size.
This creates a lot of wasted space and memory, not to mention pushes medium length strings across cache line boundaries for very little benefit - the ability to directly index/slice strings without accidentally corrupting data.
Now obviously you want to avoid accidentally corrupting data, the tradeoff comes down to whether you need direct, arbitrary indexing, or if it's worth doing some processing to determine the correct place to split in order to make space gains.
The technical world has come down overwhelmingly in favour of the latter, and that's why you see hardly anyone using utf-32. It's simply not as good a solution for most real world concerns.
Now with UTF-16, the "normal" characters are the ones in the basic multilingual plane that fit in a single UTF-16 code point.
UTF-16 has its own warts, but invalid code units and non-shortest forms are exclusive to UTF-8.
[1] http://www.sans.org/security-resources/malwarefaq/wnt-unicod...
[2] http://cve.mitre.org/cgi-bin/cvename.cgi?name=CVE-2008-2938
https://w3techs.com/technologies/history_overview/character_...
UTF16 isn’t good enough for web: even for a content in Ukrainian or Hebrew languages, UTF8 saves a sizeable bandwidth because spaces, punctuation marks, newlines, digits, English-inspired HTML tags — in UTF8 they all encode in 1 byte per character, and for the web, bandwidth matters.
Am I reading that site incorrectly? It says: UTF-8: 87.2%, not unicode.
Then down below:
" The following character encodings are used by less than 0.1% of the websites"
UTF-16
https://w3techs.com/technologies/overview/character_encoding...
I think UTF8 should be the default and only format for storing text attributes in all databases and all other text encodings should be removed from database systems.
Annoying as it is to deal with, our history as computer scientists demands that we maintain compatibility with older systems and encoding formats that were once used but are now almost forgotten. If we removed all the other encoding formats (code paths that, while underused, still function perfectly fine) we would lose the ability to parse and manipulate a lot of old data.
I work with multilingual text processing applications, and I strongly support that concept. A guideline of "use UTF8 or die" works well and avoids lots of headaches - it is the most efficient encoding for in-memory use (unless you work mostly with Asian charsets where UTF16 has a size advantage) and it is compatible with all legal data, so it's quite effective to have a policy that 100% of your functions/API/datastructures/databases pass only UTF8 data, and when other encodings are needed (e.g. file import/export) then at the very edge of your application the data is converted to that something else.
Having a mix of encodings is a time bomb that sooner or later blows up as nasty bugs.
Sometimes vector of objects that include other information, like glyphs etc.
IMHO, UTF-16 is the worst of both worlds. It breaks backwards compatibility in the simple case and wastes storage, but still has to have complex multi-byte decoding because it's not a fixed length encoding.
UTF-8 is probably the best compromise of the lot, with the advantages of UTF-32 being outweighed by the massive overhead in the most common case.
Keeping the backwards-compatibility heuristic the same makes sense.
Sites that want UTF-8 can ask for it.
This is just my first thought. Seems that the job of ICU is transfered to the OS or web browser.
Languages are horribly complicated. The Turkish ı/İ issue makes capitalization a locale-dependent thing, and things like German ß/ẞ/ss/SS make case conversion in general mind-boggling. The treatment of diacritics in Latin script for collation purposes differs very heavily between major European languages, so sorting and searching are again locale-dependent. And by the time you're dealing with the locale mess of languages, handling locale-specific number, date, and time representations is pretty much trivial.
The need for giant Unicode character tables and CLDR tables, or tables that capture similar information, is quite frankly necessary to handle internationalization to any substantial degree.
It's worse. Sorting is dependent on the task at hand. http://userguide.icu-project.org/collation: "For example, in German dictionaries, "öf" would come before "of". In phone books the situation is the exact opposite."
That page has lots more 'interesting' cases, for example:
"Some French dictionary ordering traditions sort accents in backwards order, from the end of the string. For example, the word "côte" sorts before "coté" because the acute accent on the final "e" is more significant than the circumflex on the "o"."
That means that, given two strings s and t such that s sorts before t, you can append characters to t to get u which sorts before s. EDIT (after reading the reply of kelnage): _for some strings s and t_
It could happen in a century or two, actually, we are seeing some language trends that do favor internationalization and simplification over localization and keeping with linguistic tradition.
I know you're not necessarily advocating it, but if our cultures change to adapt to our technological limitations, that's the reverse of what I think should be happening - there's a problem with the tech.
Wrong: up to 4 bytes UTF16, and up to 6 bytes UTF8.
> Cyrillic, Hebrew and several other popular Unicode blocks are 2 bytes both in UTF-16 and UTF-8.
Cyrillic, Hebrew and several other languages still have spaces and punctuation, that take a single byte in UTF8. Now it’s 2016, RAM and storage are cheap and declining, but CPU branch misprediction cost is same 20 cycles and not going to decline.
> plain Windows edit control (until Vista)
Windows XP is 14 years old, and now in 2016 it’s market share is less then 3%. Who cares what was before Vista?
> In C++, there is no way to return Unicode from std::exception::what() other than using UTF-8.
The exception that are part of STL don’t return Unicode at all, they are in English.
If you throw your custom exceptions, return non-English messages in exception::what() in utf-8, catch std::exception and call what() — you’ll get English error messages for STL-thrown exceptions, and non-English error messages for your custom exceptions.
I’m not sure mixing GUI languages in a single app is always a right thing.
> First, the application must be compiled as Unicode-aware
The oldest visual studio I have installed is 2008 (because I sometimes develop for WinCE). I’ve just created a new C++ console application project, and by default it already Unicode-aware.
So, for anyone using Microsoft IDE, this requirement is not a problem.
I haven't checked your other claims but this stands out:
> The exception that are part of STL don’t return Unicode at all, they are in English.
Do you mean they return the text as bytes using some (likely ASCII) character encoding and all the text characters are in ASCII range?
There Ain't No Such Thing As Plain Text. (2003) http://www.joelonsoftware.com/articles/Unicode.html
If you rely on std::exception::what() while building a localizable software, you’ll end with inconsistent GUI language. Because some exceptions (that are part of STL) will return English messages, other exceptions (that aren’t part of STL) will return non-English messages.
This means if you’re developing anything localizable, you can’t rely on std::exception::what().
Then why care about it’s prototype?
Why care about it's prototype? You may want to embed into what() unicode strings that describe the error and came from elsewhere. E.g. a path, a URL, an XML element id, etc. from the context the exception originated. It may be shown to the user or written to the log. Localization is irrelevant here.
Yes, I love that every byte transmitted on the Internet still reserves code points for controlling teletype (or similar) machines.
I remember a project (circa 1999) I worked on which was a feature phone HTML 3.4 browser and email client (one of the first). The browser/ip stack handled only ascii/code page characters to begin with. To my surprise it was decided to encode text on the platform using utf-16. Thus the entire code base was converted to use 16 bit code points (UCS-2). On a resource constrained platform (~300k ram IFIRC), better, I think, would have been update the renderer and email client to understand utf8.
Nice as it might be to have the idea that utf16, or utf32 were a "character" it is as has been pointed out not the case, and when you look into language you can see how it never can be that simple.
As the trade-off, directly indexing into strings is... Either not possible or discouraged, and often relies on an opaque(?) indexing class.
The main weirdness I have encountered so far is that the Regex functions operate only on the old, objective-c method of indexing, so a little swizzling is required to handle things properly.
You can't really "disable utf-8" on Linux. You can change how things are encoded when displaying or saving. (via locale/lang variables) But if the app wants to create a file named "0xE2 0x98 0x83" (binary version of course), it's still free to do that.
You could probably write some filter using fusefs, but in practice... I think you should configure the servers / clients to agree on encoding instead. Better supported and shouldn't be that much work.
But it won't allow you to force ISO-8859-1 in this form. However you could filter out non-ASCII characters.
Encoded to UTF-8 it becomes EF BB BF. Encoding to UTF-16 big endian it will become FE FF. Encoding it to UTF-16 little endian it becomes FF FE.
Converting it back from UTF-8 always gives you U+FEFF since UTF-8 doesn't care about endianess. Converting it back from UTF-16 using the correct endianess gives you U+FEFF. Converting it using the wrong endianess gives you U+FFFE which is defined by unicode as a "non character" that should never appear in text.
Note that the "BOM" in this case means storing the U+FEFF character in UTF-8 form (just as UTF-16 stores it in the appropriate endianness). This means that the result would be EF BB BF.
The BOM allows to distinguish a byte stream between non-Unicode, UTF8, UTF16 and UTF32.
Like it or not, but it's part of the standard:
Yes, it can be used to distinguish a UTF-8 stream but it's not recommended. One issue is you can't tell if the BOM is not valid text in some other non-unicode encoding.
I'm curious where you've encountered missing content-encoding headers or other OOB indicators where it wasn't because of programmer error or laziness.
If a specification says “something may be encountered”, for me, when I write my software, it means I must support that thing. Otherwise, the software won’t conform to the spec.
> I'm curious where you've encountered missing content-encoding headers or other OOB indicators where it wasn't because of programmer error or laziness.
Everywhere.
Most filesystems don’t have encoding headers for their text files. Most databases don’t have headers for their blob columns.
Only web that has encoding headers.
I misread what you wrote about where you saw no idication that it was UTF-8. You were talking about places other than the web.
BOM for UTF-8 text files seems to be a Microsoft thing. Everyone else just defaults to UTF-8. But you can't be sure that it's a UTF-8 BOM or some other encoding. Most editors let the user overide what it is.
Why would you store text in a blob column? If a database can't handle UTF-8 in it's text column it needs to be fixed (or taken out back an shot).
I’m a Windows developer. In my world, a program should generate its output in whatever format user wants it to be.
When I press “File/Save as” in visual studio and click on the down arrow icon, I see a choice of more than 100 different encodings (including all flavors of Unicode with and without the BOM), and independent choice of 3 line endings (Window, mac, Unix).
> BOM for UTF-8 text files seems to be a Microsoft thing
Practically — maybe, most Microsoft apps tend to understand those BOMs, and most *nix tools don’t, even on input.
Officially — definitely no, we both saw the spec on unicode.org.
When generating output for a user, letting them choose is a good idea. But for interop with other programs I leave it off unless the program needs it.
> Officially — definitely no, we both saw the spec on unicode.org.
The spec says the BOM is optional. Some Microsoft programs however require it.
Plain text isn’t exactly a machine-friendly format.
If you want to interop with other programs, the good choice is e.g. XML. That has this encoding problem fixed as a part of the standard.
> The spec says the BOM is optional. Some Microsoft programs however require it.
Could you please name a Microsoft program that you think requires a BOM?
I’m asking because I have completely different experience. For me, Microsoft programs open text files just fine, with or without the BOM. But most *nix and osx programs show me garbage instead of BOM.
Works fine for unix. :D
> Could you please name a Microsoft program that you think requires a BOM?
Visual C++ off the top of my head. It mangles UTF-8 string literals without the BOM in the source code.
> For me, Microsoft programs open text files just fine, with or without the BOM. But most *nix and osx programs show me garbage instead of BOM.
That's what I was trying to say about the BOM being prevalent on the Windows side of the fence. Some programs require it, some always generate it so most program now accept it.
On the unix/osx side everyone switched to UTF-8 so the BOM is redundant. Everything is UTF-8 so the silliness of this needs a BOM that doesn't need a BOM doesn't exist. Good example of what the "UTF-8 Everywhere" site is trying to promote.
Personally I really wish Microsoft would eventually fix their UTF-8 codepage. Would be so nice not having to convert to/from UTF-16 at the Win32 API boundary.
The trend towards higher-level data formats is universal across all OSes.
Even on Unix, users typically read html, write odf or docx both being xml, print PostScript, etc.
Plain text is friendly towards developers. But it’s neither interop-friendly nor user friendly.
> Visual C++ off the top of my head
Only the C++ compiler. MS can’t change the compiler because backward compatibility. The IDE however works fine with such files.
> "UTF-8 Everywhere" site is trying to promote.
The transition is going to be expensive, because most languages and frameworks (C++/MFC/ATL/QT, .NET languages, JVM languages, Python, etc) use Unicode (USC2 or UT16) strings for decades already.
To justify the costs, the benefits of the transition must be substantial.
And there aren’t any.
Kind of got off track here. You can process a lot of formats as text (html, css, xml, etc). So a BOM there is unnecessary and sometimes detrimental. On the unix side there are a lot of text utilities that do useful things that you can do on these formats. That's probably why BOMs are non existent there.
> MS can’t change the compiler because backward compatibility.
You care to tell MS that? Every single time I've done a major VS upgrade my code had to be changed because something that was valid before stopped being valid.
> And there aren’t any.
If you can't see any benefit of using UTF-8 then I'm done debating with you.