Is the use of “utf8=✓” preferable to “utf8=true”?
programmers.stackexchange.com
programmers.stackexchange.com
After while Programmers was becoming a bit too much of a dumping ground so tightened up the rules to make it less random and chatty and more of a valid site in it's own right for topics around software development but which are not directly programming related.
Stackoverflow (Objective): Why does this code give me a syntax error?
Programmers (Subjective): What programming methodology best fits my project and team?
A question like this tends to elicit a much more forceful response on P.SE than it will on Stack Overflow. Largely because the folks who moderate P.SE aren't terribly fond of the common perception that their site is just a dumping ground for questions that are too wishy-washy for SO.
Quora couldn't grow because they wanted to become THE site, SO is smart enough to pander to niches.
Stackoverflow is for stuff strictly about code, APIs, languages, syntax, etc.
Hmmm... there's something wrong with that idea, but I can't quite put my finger on it....
This isn't a config file, this is the query string of a URL, or more importantly the POST data of a form.
From the article:
> By default, older versions of IE (<=8) will submit form data in Latin-1 encoding if possible. By including a character that can't be expressed in Latin-1, IE is forced to use UTF-8 encoding for its form submissions, which simplifies various backend processes, for example database persistence.
> If the parameter was instead utf8=true then this wouldn't trigger the UTF-8 encoding in these browsers
I think your missing the sarcasm...
Wait, true...true
X is used variously as true and false in Different languages, so why would you assume it means false in a language you have never used?
˙ɔıuoɹı sı ʇɹoddns ǝpoɔıun ou s,ǝɹǝɥʇ ʇɐɥʇ ǝʇɐɔıpuı oʇ ʇı buısn os 'ɹǝʇɔɐɹɐɥɔ ǝpoɔıun ɐ ɟןǝsʇı uı sı ssoɹɔ ʇɐɥʇ
All characters -- including every one in this post -- are unicode characters. The definingly-useful characteristic of the cross (or checkmark) is that it can't be represented in latin1.
Edit: Please :-)
which has this javascript code to convert:
SampleUtf8Char= would be more clear.
<!--[if lt IE 8]><input name="utf8" type="hidden" value="✓" /><![endif]-->
Has anyone experimented with doing that?
EDIT: found the commit with some git log magic - https://github.com/rails/rails/commit/c616089
edited to add: This seems related to a problem various "try to sound like English" programming languages (e.g. Inform) have, where it is easy to assume invalid syntax will be valid because it's valid English.
What? This is wrong. UTF-8 encodes a lot more than just ASCII.
UTF-8 is compatible with ASCII in that all of the characters ASCII and Unicode have in common are represented the same way in ASCII and UTF-8. Going beyond ASCII involves the introduction of multi-byte representations in UTF-8, and that takes you smoothly (that is, no surrogate pairs) out into the entire rest of Unicode. As a bonus, it's always possible to verify that a given string of bytes is valid UTF-8, given that there is a nontrivial structure imposed on UTF-8 multi-byte encodings that is very unlikely to occur by chance in any non-UTF-8 sequence of bytes.
It gets far worse in 3-byte UTF8 characters, but I don't believe any of them exist natively in Latin1 (see: euro symbol)
Assuming I'm reading these various character tables right, at least ;)
So a more accurate version of what you quoted would be "UTF-8 and Latin-1 only overlap for 7-bit ASCII"
0xC2A2 will be rendered as ¢ only if it's encoded in UTF-16/UCS-2 big endian and misinterpreted as ISO-8859-1/Windows-1252.
If it's encoded in little endian (much more common on Intel x86 computers), then it would be rendered as ¢Â when misinterpreted.
At any rate, if you were to encode ¢ in UTF-16BE, it would be 0x00a2, not 0xc2a2. If a piece of software then misinterpreted it as latin1, likely you'd get nothing at all due to the embedded NUL.
$ echo -n ¢ | iconv -f UTF-8 -t UTF-16BE | hexdump -C
00000000 00 a2 |..|Indeed. I either completely misread the parent post, or else it said something different when I responded to it (knowing myself, I'm going with the former).
The point of the checkmark, therefore, is to put it in a hidden form field. That way, no matter what the user types, there will still be at least one non-Latin-1 character. That will force IE to use UTF-8, and you can check to make sure this actually happened by checking the value of the form field: if it's not set correctly, then you know there may be trouble.
That's correct. If you send your HTML document with a charset of UTF-8 (In the Content-Type header) then IE will submit forms using UTF-8 even if the user doesn't input any UTF-8 characters. Unless the user changes the encoding, but I have yet to hear a compelling reason why an ordinary user would do that under ordinary circumstances.
> The snowman hack serves to prevent the corruption from spreading.
It's clever, but the framework could also just reject POST and GET requests which contain invalid UTF-8 characters. (I'm flabbergasted that Ruby doesn't do this[1].) Otherwise a malicious user could try to inject non-UTF-8 characters into your database by sending crafted requests which nevertheless contain the "utf8=✓". And speaking from experience, you do not want to have to deal with encoding problems in your database.
[1] http://stackoverflow.com/questions/3222013/what-is-the-snowm...