The point of the checkmark, therefore, is to put it in a hidden form field. That way, no matter what the user types, there will still be at least one non-Latin-1 character. That will force IE to use UTF-8, and you can check to make sure this actually happened by checking the value of the form field: if it's not set correctly, then you know there may be trouble.
That's correct. If you send your HTML document with a charset of UTF-8 (In the Content-Type header) then IE will submit forms using UTF-8 even if the user doesn't input any UTF-8 characters. Unless the user changes the encoding, but I have yet to hear a compelling reason why an ordinary user would do that under ordinary circumstances.
> The snowman hack serves to prevent the corruption from spreading.
It's clever, but the framework could also just reject POST and GET requests which contain invalid UTF-8 characters. (I'm flabbergasted that Ruby doesn't do this[1].) Otherwise a malicious user could try to inject non-UTF-8 characters into your database by sending crafted requests which nevertheless contain the "utf8=✓". And speaking from experience, you do not want to have to deal with encoding problems in your database.
[1] http://stackoverflow.com/questions/3222013/what-is-the-snowm...
What? This is wrong. UTF-8 encodes a lot more than just ASCII.
UTF-8 is compatible with ASCII in that all of the characters ASCII and Unicode have in common are represented the same way in ASCII and UTF-8. Going beyond ASCII involves the introduction of multi-byte representations in UTF-8, and that takes you smoothly (that is, no surrogate pairs) out into the entire rest of Unicode. As a bonus, it's always possible to verify that a given string of bytes is valid UTF-8, given that there is a nontrivial structure imposed on UTF-8 multi-byte encodings that is very unlikely to occur by chance in any non-UTF-8 sequence of bytes.
It gets far worse in 3-byte UTF8 characters, but I don't believe any of them exist natively in Latin1 (see: euro symbol)
Assuming I'm reading these various character tables right, at least ;)
So a more accurate version of what you quoted would be "UTF-8 and Latin-1 only overlap for 7-bit ASCII"
0xC2A2 will be rendered as ¢ only if it's encoded in UTF-16/UCS-2 big endian and misinterpreted as ISO-8859-1/Windows-1252.
If it's encoded in little endian (much more common on Intel x86 computers), then it would be rendered as ¢Â when misinterpreted.
At any rate, if you were to encode ¢ in UTF-16BE, it would be 0x00a2, not 0xc2a2. If a piece of software then misinterpreted it as latin1, likely you'd get nothing at all due to the embedded NUL.
$ echo -n ¢ | iconv -f UTF-8 -t UTF-16BE | hexdump -C
00000000 00 a2 |..|Indeed. I either completely misread the parent post, or else it said something different when I responded to it (knowing myself, I'm going with the former).