Unicode over 60 percent of the web
googleblog.blogspot.com
googleblog.blogspot.com
UNICODE is a good thing because it provides a codepoint for every character that we care about, instead of having a 256-character subset for every groups of languages and needing complicated software to puzzle out how to convert from one subset to the other. Unicode allows fantastic stuff such as upper/lowercasing text including all the weird letters that you previously had to special-case.
ASCII used to be a good thing because it allowed people to ship around basic English and Cobol code without any worries, but is actually pretty evil because people from Anglosaxon countries assume that every other bit of text is composed of English and Cobol.
Having a notion of ENCODINGS is useful if you occasionally get bits of text that are neither English nor Cobol. You still needed different encodings for different groups of languages, and arcane mechanisms to provide hints on which encoding is meant, at least if you got non-English bits of text. The very notion of an Encoding scares the people who used to think the world consists of English and Cobol.
UTF-8 is a very reasonable encoding that can be used to represent all of Unicode while being Ascii-compatible. Hence it is a sane choice as a default encoding for people who are scared of having to think about encodings. Because UTF-8 is not the only encoding out there, Unicode-compatible programs accept Unicode text in many other encodings, including those that cannot represent the full range of Unicode and are only a good choice for some people but not others.
tl;dr: non-UTF-8 text can (and should) still be read as unicode codepoints. Ignoring the >=40% of texts out there or saying that they're "not Unicode" doesn't help anybody.
From a security perspective, if you don't include the right encoding meta-data, an attacker can include an XSS attack in UTF-7. Your server reads normal (but gibberish) ASCII, escapes it properly, then shleps it to the client who thinks it's UTF-7 (because of all the UTF-7 chars) and suddenly they are running malicious javascript that you didn't escape.
If you're scraping sites, the one thing more annoying than trying to guess the encoding, and that's dealing with multiple encodings in the same document.
Putting UTF-8 characters in a document doesn't mean you are using "Unicode". It means you are using some undeclared encoding, which is not a step forwards.
User agents must not support the CESU-8, UTF-7, BOCU-1 and SCSU encodings.
and major browsers removed support, e.g.:
Are you saying that we should still use character encodings rather than UTF encodings, or are you saying that we shouldn't assume that raw text is ASCII, or are you saying something else?
Unicode is, essentially, nothing unless it is encoded. When you encode it you must decide whether to use one byte, two bytes, three bytes or four bytes.
Only UTF-8 and UTF-32 are really big enough to hold the world's characters. everything else is a fudge.
ASCII was never Anglo-Saxon: it was always American. COBOL is a red herring here, too. ASCII was all about teleprinters and was a clever use of 8-Bits for its time. Bell labs invented ASCII and Bell labs invented UTF-8.
UTF-8 is much better than reasonable. It is a compact way to represent Unicode while preventing the western world from having to re-encode every text document. That's a lot of good news the internet.
Are you saying that we shouldn't ignore other charsets as they are still valid Unicode? If so I agree up to a point, the point being that there is no longer any need to have any other unicode encoding other than UTF-8. If you need to access your local characters as a byte array: choose your internal encoding and translate, do your magic, and then spit out UTF-8 again then we can all simply read the same documents without the need for over-complexity.
So is GB18030.
Edit: and UTF-16, of course (just don't confuse it with UCS-2)
But it does beg the question why would you use UTF-16? Yes if you have more than 512 characters in your common script then OK I can see it might make sense, but not much. UTF-8 will still average out in a reasonable way.
Unicode is not a text encoding, it's a standard that assigns a number to every character.Saying that a particular text is "Unicode" doesn't give any info on how to decode it, we should just say that a given text is UTF-7,8,16,32, etc.
Now, consider that if you decide to use utf-16 to publish your Japanese text as an HTML document (whose syntax is entirely ascii based), you might lose its "saving advantage" and could easily end up with text that is actually larger in utf-16 than in utf-8.
Also take a look at this: http://programmers.stackexchange.com/questions/102205/should...
Example: The Korean text of the Universal Declaration of Human Rights [1] is 8.1KB in EUC-KR and 11.2KB in UTF-8. When compressed with bzip2, it's only 3.1KB and 3.2KB, respectively. I assume Japanese would behave similarly.
[1] http://www.ohchr.org/EN/UDHR/Pages/Language.aspx?LangID=kkn
So basically it's not fixed width in any meaningful way.
There is no such thing as "UTF-32 character".
Abstract character is not code point.
UTF-32 _is_ fixed width because it's defined on code points, not glyphs, not characters.
Mojibake (文字化け?) (IPA: [modʑibake]; lit. "unintelligible sequence of characters"), from the Japanese 文字 (moji) "character" + 化け (bake) "change", is the occurrence of incorrect, unreadable characters shown when software fails to render text correctly according to its associated character encoding.
It's certainly very easy to do - I had the problem a little while back setting up a little web app where the web server and MySQL database were both UTF-8, but the db connection was defaulting to ISO-8859-1 or something, causing all sorts of issues with curly quotes etc.
And it helps that almost all editors, these days, also default to UTF-8 encoding. So you can just copy/paste special characters and it works...
I think they are measuring the encoding that is picked before adding a page to the index. Which would be a combination of explicit information (headers, xml and html metadata) and heuristics (which may fail, but are useful when the explicit information is missing or — even though ignoring explicit metadata is bad — obviously incorrect).
It also seems they are looking for a subset encoding once they have the metadata; the posts describe explicitly labelling ASCII when the contents are within that subset.
But their methodology has changed from the previous two posts: latin above ascii in 2008 didn't exist in http://googleblog.blogspot.com/2010/01/unicode-nearing-50-of... or http://googleblog.blogspot.com/2008/05/moving-to-unicode-51.... .
http://stackoverflow.com/questions/3616359
Also the encoding for .properties files is specified as Latin-1
Here's my take: UTF-8 should, rightly, be the only interchangeable text format of choice for the right-minded individual. The other UTFs should be internal formats used in-memory or on disk cache / database / etc
Why? Simple: UTF-8 rocks!
It's a beautifully designed, backwardly compatible, nifty piece of back-of-a-napkin genius. Simple to code and decode (once you understand it), simple to check with a regular expression (once you realise it is essentially just a token with a set number of chars), and simple to add the wealth of the world's characters to your app with a reasonable amount of code. Plus, it's compact in the way Huffman coding is compact (at least from a western perspective).
Also, no-one should ever be using (char ❄) in the 21st century unless it is to temporarily hold a UTF-8 string before converting to wchar.
Why? Because you almost certainly don't have a (char ❄), you have a UTF-8 sequence (^^^see above). Sadly this makes your memory mapped files slightly redundant. But don't fret, this is the future: convert them to 32-bits and release them. Be happy that you can now treat any character sequence, in common usage, in the whole of humanity like an array.
As for UTF-16, why bother? It's neither compact, clever, nor big enough to hold every character on the internet. 💩 needs more than 16 bits and everyone, now, needs to support a poop with eyes.
tl;dr: Share UTF-8 promiscuously, keep UTF-32 for private moments. Don't dally with UTF-16, she's an old tease and can't handle poop.
note: ❄ = asterisk :)
There may be a distinct size advantage for some asian cultures to using UTF-16 instead of UTF-8 as it will allow for encoding more of the glyphs without having to add more overhead bits. How much this saves in reality I'm not sure.
The advantage of moving to one standard, plus the backwards compatibility for most of the internet will outweigh any small size advantage that UTF-16 will have on real file sizes for some countries / cultures.
A saving of ⅓ on a text file will barely be noticed in a world seemingly governed by Moore's law in nearly every future metric.
(also my bad for using UTF-16 where I meant 16-bit unicode character arrays :))
If you develop an application for the international market you should probably go with UTF-8 as well. Maybe if you develop only for the Asian market it's a bit different. But in my experience significant amounts of text usually come in some form of data or markup format (HTML, XML, JSON, etc.) and usually those markup formats are defined in the ASCII subset. So UTF-8 still wins. Just take a random page from the Japanese Wikipedia and encode it in UTF-8 and in UTF-16. You'll see that UTF-8 almost always wins.
So his inventions of Unix and UTF-8 now dominate both the back- and front-ends of the internet.
...which indicates that nearly all unicode is UTF-8.
"Commonly used character encodings on the Web include ISO-8859-1 (also referred to as "Latin-1"; usable for most Western European languages), ISO-8859-5 ..."
KOI8-R was, then Windows-1251. Now it's often UTF-8.
It actually doesn't because of all the HTML tags being ASCII.
What I really want to know is the breakdown by alphabet. How many sites are in Cyrillic? Kanji? Arabic? Thai?
Actually, outside of listings of all unicode/code pages, I'm 98% sure there is a vast quantity of unicode characters that is not used once (organically) on the entire Internet.
I'd bet even money you can craft a two-character 'word' such that you would be the top Google result for it if you use it just once or twice in any context on any page Google indexes, just because you're the first person to use those characters organically, to say nothing of together. /s
> The more documents that are in Unicode, the less likely you will see mangled characters (what Japanese call mojibake) when you’re surfing the web.
Hmm. Well, I know that's true in theory, but my personal experience is that the occurrence of mojibake has increased lately. I even wrote about it here on HN: http://news.ycombinator.com/item?id=2075010
HTML is defined to use Unicode as the document character set. But the charaters can be represented as byte-streams using different encodings, UTF-8 beeing one encoding, ISO-8859-1 beeing another encoding.
> The "charset" parameter identifies a character encoding, which is a method of converting a sequence of bytes into a sequence of characters.
A lot of people seem to confuse Unicode with the UTF-encodings.
That document is badly-written; a more reasonable way to interpret it is to conclude that user agents will use some form of Unicode internally, after converting whatever character encoding the document they received used. Which is, indeed, a very reasonable way to design your software, but it doesn't make Latin-1 (for example) a Unicode encoding by any reasonable standard.
> A lot of people seem to confuse Unicode with the UTF-encodings.
True. I do not.