Unicode shenanigans: Martine écrit en UTF-8
blog.poisson.chat
blog.poisson.chat
"what Planet software is Planet Haskell using that it's doing these odd things - oh, it's Venus? https://github.com/rubys/venus ... wow that's old... and doesn't that use something like filesystem storage for the feeds and couldn't this happen if you stored the xml with no character set specified and the parser messed it up on the way back in? Which since it's ending up with a windows encoding...wait, why am I remembering any of this....oh. checks credits for Venus, recordscratch yes, that's me, wondering why I fixed venus to run on windows 18 years ago"
https://github.com/rubys/venus/commit/210781768705d20dc3cbe6...
Sorry about that. As I recall at the time I was the only person using Venus on windows, and I was just running a hacked up version on a local machine. There's some conditionals in there about whether it uses libxml2 or not (in modern python, it wouldn't need to), that call doesn't take a charset parameter and my guess is the problems begin as soon as libxml2 tries to parse the file on disk. I think my own version was frankensteined to the extent of using a sqlite back end so I didn't have to deal with windows files any more.
The author of Venus, Sam Ruby, is on here (rubys) but it looks like he hasn't checked in for a long time; last I saw he was over at fly.io.
Oh and the even funnier part of this all is, back when this was written, Sam's blog was THE go to place to look up the list of Mojibake mistakes you'd made...
https://intertwingly.net/stories/2004/04/14/i18n.html#Cleani...
And I fully agree with footnote 2, why is an extended Unicode character being used in place of the apostrophe? https://tedclancy.wordpress.com/2015/06/03/which-unicode-cha...
Some reasons:
> The Unicode character ’ (U+2019 right single quotation mark) is used for both a typographic apostrophe and a single right (closing) quotation mark.[1] This is due to the many fonts and character sets (such as CP1252) that unified the characters into a single code point, and the difficulty of software distinguishing which character is intended by a user's typing.[2] There are arguments that the typographic apostrophe should be a different code point, U+02BC modifier letter apostrophe.[3][better source needed]
> The straight apostrophe ' (the "ASCII apostrophe", U+0027 ' apostrophe) is even more ambiguous, as it could also be intended as a left or right quotation mark, or a prime symbol.
Actually I ran into a separate issue on Windows where Python will automatically replace the line endings depending on the OS. So I had to specify newline='\n' as an argument to open() or it would alter the newlines to Windows format in the output.
(My fault for not running it in WSL, I guess.)
Unicode has only one apostrophe. The same one as ASCII apostrophe. The problem here is people using right single qoute as a "fancy" apostrophe when they should be using a font that renders apostrophe in the desired way. I have to fix this junk all the time in CD-text and Musicbrainz metadata.
Second: no, Unicode definitely considers other characters to be apostrophes:
>>> unicodedata.name(chr(700)) # https://en.wikipedia.org/wiki/Modifier_letter_apostrophe
'MODIFIER LETTER APOSTROPHE'
>>> unicodedata.name(chr(1370)) # https://en.wikipedia.org/wiki/Armenian_(Unicode_block)
'ARMENIAN APOSTROPHE'
>>> unicodedata.name(chr(2036)) # https://en.wikipedia.org/wiki/N%27Ko_script
'NKO HIGH TONE APOSTROPHE'
>>> unicodedata.name(chr(2037))
'NKO LOW TONE APOSTROPHE'
>>> unicodedata.name(chr(65287)) # https://en.wikipedia.org/wiki/Halfwidth_and_fullwidth_forms
'FULLWIDTH APOSTROPHE'
>>> unicodedata.name(chr(917543)) # https://en.wikipedia.org/wiki/Tags_(Unicode_block)
'TAG APOSTROPHE'
There's even a "double apostrophe" which is definitely not any kind of quotation mark, semantically (https://en.wikipedia.org/wiki/Modifier_letter_double_apostro...) .For that matter, there's a "Greek question mark" (https://en.wikipedia.org/wiki/Question_mark#Greek_question_m...) which is semantically equivalent to ? for Greek text, but rendered identically to ; in most fonts. And there are some CJK ideographs (https://en.wikipedia.org/wiki/Ghost_characters) which represent completely fake characters not used in any writing (including in any other East Asian language), which were included in an old Japanese character standard by mistake and then copied into Unicode.
"Unicode only has one" is the start of a lot of false claims. It is very much "designed by committee".
Second: no, there is only a single apostrophe fit for this context, the ones are used for different purposes
- Unicode has only one apostrophe (the ASCII one literally named that)
When Unicode says
- 2019 is preferred for apostrophe
You can’t appeal to Unicode and then ignore them in the next breath.
I was responding to the list of symbols named "apostrophe", where zahlman also seems to follow the consistent logic of only listing apostrophes, not quotes
You: No, that’s completely true what they said there
You: Wait, you should take that up with them…
Just don’t speak out of your mouth with “Unicode” when you aren’t prepared to get it thrown back in your face? Doesn’t seem difficult.
> I was responding to the list of symbols named "apostrophe", where zahlman also seems to follow the consistent logic of only listing apostrophes, not quotes
Oh wow they’re named apostrophe? How great. I’ll start using this “start of header” character in my emails, that is probably so fit for purpose.
It is true what they said within the context of the conversation where the standard notes are either explicitly or implicitly rejected.
> that is probably so fit for purpose.
Or probably not. Unlike in this case where apostrophe is definitely fit for purpose
No, the ASCII ‘quotes’ are inch and foot markers. Relying on a font to render the the inch and foot markers as quotes is a … unique … approach.
> I have to fix this junk all the time in CD-text and Musicbrainz metadata.
Wait, are you the guy whose Musicbrainz titles are constantly replacing my carefully- and properly-punctuated ones every time I sync? I beg you to reconsider.
There are special symbols for those. There is no other symbol for the apostrophe
> Relying on a font to render the the inch and foot markers as quotes is a … unique … approach
No, you would rely on a font to render those dedicated inch symbols as inch markers
Besides you can only do the display substitution with context awareness, so that ' as inches won't be changed to quotes, but ' in titles will be
ibus. In my X settings the physical Caps Lock key is turned into Compose, and then I can easily type ‘ with Compose <' (and vice-versa), ’ with Compose >' (ditto), “ with Compose <" (ditto), ” with Compose >" (ditto), ß with Compose ss, þ with Compose th, Þ with Compose TH, æ with Compose ae, Æ with Compose AE, … with Compose .., — with Compose ---, – with Compose --. and so forth.
https://github.com/kragen/xcompose offers an excellent XCompose file with over a thousand wonderful Compose mappings. Highly recommended!
0027 APOSTROPHE
[snip]
* 2019 is preferred for apostrophe
https://unicode.org/Public/UNIDATA/NamesList.txt(Or even the dreaded case of using the Philippine Peso sign (which is simply named in Unicode as the PESO SIGN) for other pesos, which is sometimes encountered in some apps! At least this is outright wrong that it is corrected immediately.)
Two code points have "apostrophe" in their name, U+0027 and U+02BC:
* https://en.wikipedia.org/wiki/Apostrophe#Unicode
There is also a third which is often used, U+2019:
> The Unicode character ’ (U+2019 right single quotation mark) is used for both a typographic apostrophe and a single right (closing) quotation mark.[1] This is due to the many fonts and character sets (such as CP1252) that unified the characters into a single code point, and the difficulty of software distinguishing which character is intended by a user's typing.[2] There are arguments that the typographic apostrophe should be a different code point, U+02BC modifier letter apostrophe.[3][better source needed]
> The straight apostrophe ' (the "ASCII apostrophe", U+0027 ' apostrophe) is even more ambiguous, as it could also be intended as a left or right quotation mark, or a prime symbol.
IIRC it's combination of some new_API_2.final() Win32 API and compilation options for C++, C#, and Java. Microsoft briefly tried to switch system locale to UTF-8 on Windows 10, but I think they've since given up on it.
Other platforms such as Electron, Android, iOS shouldn't have this problem; those should be UTF-8 native.
"Martine écrit en UTF-8" is what you get if you interpret the UTF-8 string "Martine écrit en UTF-8" (Martine writes in UTF-8) as Latin-1. It's not as common as it once was, but that kind of encoding issue used to be encountered fairly often by French speakers on the Web.
https://www.casterman.com/Jeunesse/Catalogue/vive-la-rentree...
EDIT: for a short course in contemporary french slang, one could do worse than https://www.google.com/search?q=martine+parodie&udm=2
(my favourite being "Martine and her parodies")