For new APIs in which legacy interoperability isn't needed, I completely approve of this document.
For new APIs in which legacy interoperability isn't needed, I completely approve of this document.
The Windows API calles UTF-16 "Unicode". Most Mac OS X APIs use UTF-16. JavaScript and Java both use UTF-16. ICU uses UTF-16. So while UTF-8 is technically superior in almost every way, it's going to be an uphill battle to standardize on it.
I appreciate that new languages like Rust and Go made the choice of UTF-8 as their native text encoding. But there's a lot of inertia for UTF-16, and I'm not sure it'll be easy to ever get free of it.
The same is true of JavaScript; while you could technically implement the strings however you want, the APIs are all oriented around UTF-16 code units. And the Windows API as well, is all built around UTF-16 code units.
The problem with all of these APIs is that they make the mistake of conflating characters and code units. They all make the assumption that a character consists of a single, fixed width integer, of some given size (16 bits in the case of UTF-16). It is better to distinguish between indexing in code units (such as bytes in UTF-8 or 16 bit integers in UTF-16) and indexing in code points, or glyphs, or whatever higher level concept you are talking about. Really, for anything higher than the code unit level, you should be dealing with variable-length strings, and not try to force that into fixed length units. With UTF-8, there's no temptation to treat a single code unit as being an independently meaningful entity, as that assumption breaks down as soon as you get past the ASCII range; while with UTF-16, it's easy to make that mistake, since it holds true for everything in the Basic Multilingual Plane, which contains most characters you're likely to encounter on a day to day basis.
var decode = function (bytes) {
return decodeURIComponent(escape(bytes));
}
var encode = function (string) {
return unescape(encodeURIComponent(string));
}
(Definitely test it before wailing about benchmarks. My guess is that whatever else you’re doing is likely much slower.)If you do care about errors, or especially if you need to deal w/ UTF-8 streams that might be chopped mid-character, use something like https://github.com/gameclosure/js.io/blob/master/packages/st...
Consider a pattern like this: A page calls document.createElement(), adds a large text node (say, the collected works of Shakespeare in text form) to it, calls window.getComputedStyle() on that element, then throws the element away. This series of DOM manipulations must go through the layout engine. If the layout engine knows only UTF-8, then the layout engine has to convert the collected works of William Shakespeare from UTF-16 to UTF-8 for no reason (as it needs an up-to-date DOM to perform CSS selector matching for the getComputedStyle() call). There is no reason to do that when it could just use UTF-16 instead and save itself the trouble.
Actually, JavaScript _does_ require UTF-16. From the ES5.1 spec:
> A conforming implementation of this Standard shall interpret characters in conformance with the Unicode Standard, Version 3.0 or later and ISO/IEC 10646-1 with either UCS-2 or UTF-16 as the adopted encoding form, implementation level 3. If the adopted ISO/IEC 10646-1 subset is not otherwise specified, it is presumed to be the BMP subset, collection 300. If the adopted encoding form is not otherwise specified, it presumed to be the UTF-16 encoding form.
'<non-BMP character>'.length == 2 '<non-BMP character>'[0] == first codepoint in the surrogate pair '<non-BMP character>'[1] == second codepoint in the surrogate pair
Any JS implementation using UTF-8 would have to convert to UTF-16 for proper answers to .length and array indexing on strings.
Alternatively, if a JavaScript implementation chose to completely ignore that particular requirement, I'd guess that approximately zero pages (outside of test cases) would break, and a few currently broken pages (that assumed sane Unicode handling) would start working.
An implementation that fails to meet the spec can be argued as not broken, even if said spec is broken.
So what? Is your goal to create useful software, or win at worthless benchmarks?
Is fast DOM manipulation important? Given that the only way for the sole scripting language on the Web to display anything or interact with the user is through DOM manipulation, I think it's worth optimizing every cycle...
If the function you're calling with UTF8 is non-trivial, converting a few dozen bytes is unlikely to make a significant difference. Benchmark it, of course, but don't be surprised if you don't need to care. Modifying the DOM is probably going to be non-trivial.
https://bugzilla.mozilla.org/show_bug.cgi?id=506431
Since then we've done things such as fast-path ASCII -> UTF-16 conversion with SSE2 instructions. Converting a few dozen bytes is unlikely to make a significant difference, but often one needs to deal with more than a few dozen bytes.
> We believe that all other encodings of Unicode (or text, in general) belong to rare edge-cases of optimization and should be avoided by mainstream users.
So if you're writing a browser that must use UTF-16 in its Javascript engine due to dumb standards... it is a reasonable performance optimizations to use UTF-16 for your strings. But how many people write Javascript engines?
There are obviously ways around this. The API could be rewritten to exchange strings less frequently or use static strings that could be replaced with handles. You can try to be clever about your buffer allocation and share one amongst all calls (but watch out for threading issues!) You could write your own allocator. But all this plumbing just increases complexity and the risk of bugs, along with adding its own performance cost.
I'm not arguing against UTF-8 as the preferred encoding for many future applications, but the "minimal overhead" example given in the manifesto isn't particularly convincing.
Would it matter in a real-world setting? I can't say for sure, because nobody I know of has tried making a production-quality UTF-8 web layout engine. But, in my mind, none of the benefits of UTF-8 (memory usage being the main one in a browser [1]) outweigh the performance risks of doing conversion. And the risk is real.
[1]: Note that you still need UTF-16 anyway, for interoperability with JavaScript. So using UTF-8 might even lead to worse memory usage, due to the necessity of duplicating strings, than a careful UTF-16-everywhere scheme that takes advantage of string buffer sharing between the JS heap and the layout engine heap would.
For legacy reasons, there probably isn't a point in changing it. But I'd be surprised if performance reasons turned out to be gating.
I wonder how long it will take until people find their balls and decide to move towards the right direction.