ASCII and Unicode quotation marks
cl.cam.ac.uk
cl.cam.ac.uk
It doesn't look that much better, and it always fucks with me at random times. That shit is on the list of annoying problems that shouldn't exist in the first place, along with the \nl\cr thing, and the txt saved as rtf thing, and the UTF-8 encoding-character-at-the-beginning-of-the-file or whatever it is called.
Someone complains your program gives them an error when they open a csv file you sent them. You tested your program, it works. You go on the phone with them for 30 minutes, try to figure out what the fuck was going on. There it is, it was opened in a program that meddles with that "" and replaces it with the “” shit.
Also, there has to be at least one time you're fucked by the "" -> “” snobbiness when you go to a random Wordpress site and paste the command they tell you to do to the command line and realize it doesn't work. You pull your hair for a couple of minutes, and there is that sneaky ” thing. Wordpress does that for anything it doesn't think as code (inb4 ”good programmers don't paste commands from wordpress to GNU+bash“).
One of the first things that I do when I set up a new Mac computer is to turn that damn "" -> “” ““““feature”””” off.
That thing's the BOM.
I'm not sure if a BOM is a good way to handle it, but saying 'just do 8-bit clean' doesn't work when you're displaying or printing the characters for humans to understand.
John Doe ‘42
An abomination ‘cause it's a damn apostrophe, not opening a quotation.
System Preferences -> Keyboard -> Text -> Use smart quotes and dashes (uncheck).
On the mac:
Option [ “ (English open quote)
Option Shift [ ” (English close quote)
A lesser-known feature is that some other quotations are also possible from that same US-English keyboard: Option \ « (French open quote)
Option Shift \ » (French close quote)
Option Shift W „ (German open quote)"This is string containing an embedded \"quoted\" string"
Then I have to think about whether or not the system I'm going to send that string to is going to "helpfully" remove the backslashes, in which case I need to write:
"This is a string containing an embedded \\"quoted\\" string"
God help you if you want to go two levels deep.
All this horrible complexity could have been avoided if we could just write:
«This is a string containing an «embedded» quoted string»
Alas.
ASCII did add <>, [], and {}, any of which could have been used for quoted strings, had the programming language designers chosen that option.
https://en.wikipedia.org/wiki/String_literal#Paired_delimite... points out that PostScript and Tcl have a string literal which allows matched quotes.
PostScript: (The quick (brown fox))
Tcl: {The quick {brown fox}}Yes, but that's a pretty rare case, much more so than embedded strings.
Even that case could be solved by having two different quotes, like Python which allows both 'string' and "string". So you could do:
«This is a string that mentions the ” character without escaping it»
“This is a string that mentions the « character without escaping it”
Yes, there are still some edge cases, like embedding both “ and « in the same string. But that's really rare.
Stop using "punctuation" when you are attempting to "delimit" text. Use a character that is not punctuation, specifically designed for "field delimiter" purposes.
Trying to do two things at once is ridiculous.
Been there. Use a lib that implements a documented standard, even a bad one. Only problem is Excel, which basically standardizes on CSV and occasionally mangles your data into malformed dates because reasons anyways.
Yes, but CSV files are record collections, they are not in 99% of cases recursive like that.
If a column contains escaped secondary documents, there's something wrong.
Erm.. text can contain tabs, too. This problem was solved so, so long ago when all the various ANSI/ASCII/whatever encodings were compiled by specifically reserving not one but two characters precisely to serve as field and record delimiters.
0x30 and 0x31 solved not only the problem of having commas or tabs in your text preventing you from treating them as field delimiters, but also allowing you to include new lines and carriage returns in your fields, too!
0x30 is the unit separator (aka field delimiter) and 0x31 is the record separator (aka, well, the record separator).
I _believe_ there was a record key on some standardized keyboard layout back in the day, too.
Edit: sorry, they are decimal, not hex. Thanks @jrochkind1
Granted, that’s the whole point, but it also makes authoring and instruction harder. (And we all know how many programmers are really just competent copy-pasters.)
I know XML et al are frustrating, but I'd rather see them than a "creative" solution. It seems like 60% of the reason we still have to deal with archaic flat formats is support for Excel.
I have spent some time working with MARC 21 binary encoding (used for library cataloging records) which uses ASCII 0x1D, 0x1E, and 0X1F as delimiters. I would def not call it appreciably more _convenient_ than a more modern 'text' record format. If it has benefits, convenience isn't really one of them.
Also most CSV parsers support quotation marks and escaping to get around the comma and new line et al problems. eg:
"full name", "address"
"Homer Simpson", "742 Evergreen Terrace,\nSpringfield"
"Bart \"El Barto\" Simpson", "742 Evergreen Terrace,\nSpringfield"
Granted it's not the prettiest and some spreadsheets really love to break the formatting upon save (cough Microsoft Excel cough) but it does work.As a side note, the best spreadsheet I've found for manipulating CSV data without breaking the formatting upon saving was OpenOffice Calc. This was a few years back before the LibreOffice fork was created as I've thankfully not needed to deal with CSV files large enough to warrant a full blown spreadsheet editor, but I would assume LibreOffice Calc would behave the same.
full name,address
Homer Simpson,"742 Evergreen Terrace,
Springfield"
"Bart ""El Barto"" Simpson","742 Evergreen Terrace,
Springfield"
(Omitted optional quotes for fields that don't need them). Quotes are escaped with "", and line breaks don't need escaping, they just have to be in a quoted field. And there is no space after a comma, except you want that space to be part of the field's value.> Omitted optional quotes for fields that don't need them
I think it's good policy to always wrap your contents in quotes regardless of whether you have a delimiter that needs quoting. And in fact many CSV marshallers will do just this.
> And there is no space after a comma
That was added purely for readability on HN. I agree it's not how you'd normally marshal the contents.
%q{This is a string with an %q{embedded quote}.}It's helpful to remember that quotes will interpret the variables inside, while apostrophes will not. Very useful for scripting the creation of scripts. Example:
"It is $time" > It is 15:22
'It is $time' > It is $time
"'$time' is $time" > '$time' is 15:22
But, yes, having single and double quoted strings is another way to avoid escaping (which Ruby and a number of other languages discussed as supporting the approach being discussed also support.)
const char * str = R"*^*(This is string containing an embedded "quoted" string)*^*";
[1] http://en.cppreference.com/w/cpp/language/string_literal say qq<I can do '" in here>;I think this is actually desirable, since in your case the escape denotes different semantics. The unescaped pairs act like quotation operators while the escaped version is a character literal.
To use formal language theory, strings containing escape characters are regular, i.e. parseable with a finite-state machine. Allowing nesting means you need a stack to find the matching pair.
Even for things where the nesting does happen during lexical analysis, it's pretty trivial to keep a count in your lexer. Lots of languages support nesting comment syntax or string interpolation, which both have equivalent difficulty.
/* nested /* comments */ don't work */ (* nested (* comments *) work *)https://wiki.dlang.org/Commenting_out_code#Nested_comments
This allows commenting out code containing comments, which can be useful when debugging or giving usage examples in the code.
To end a string, use the » character.
?«Oh, you just use the « character.»
Parse error. Unexpected EOF.
To clarify: we’d still need escaping but in fewer cases.
Only if there was no chance of unbalanced quotes to need to be in the string.
Okay.
``quoted''
Is how you're supposed to write short quotes in the TeX/LaTeX typesetting system.[edit: My point being that the author seems to think this type of quoting originated with X11... which is actually newer than TeX (X11 was first released in 1984), and that the prevalence of this type of quoting likely originated with TeX when it was released in 1978... which isn't mentioned at all in the article. In fact, since TeX/LaTeX is what all the CS, Physics, and Math types were using for journal articles, it is likely the X11 font bitmap glyphs were intentionally shaped like curly quotes to make editing your TeX source files prettier.
At least, that's how I remember it...]
\left( \right)
\left{ \right}
`` ''
Why not \left" \right"
\left' \right'
Better yet, make it completely DRY: \( \)
\{ \}
\`` \''
\` \'EDIT: Nope.
As opposed to regular dash: -
Hitting the button to the right of [0] outputs a hyphen ("-"). Holding alt/option when hitting it will output an en dash ("–"), holding both alt/option and shift will output an em dash ("—").
- = -
⌥- = –
⇧⌥- = —
Test it with:
setxkbmap -option compose:menu
Then press the menu/compose key (next to right control), then C, then =. You get €.Try compose, 1, 2 for ½.
Compose, ^, 3 for ³.
Compose, A, : for Ä.
It's pretty intuitive for the most useful characters, and easily the fastest way I have of typing the ö, ñ and å in various colleagues' names.
I don't yet speak the language of my adopted country, so it's better for me to keep []{} etc where I like them in the British layout, and use three keypresses for typing the ø in a (place) name like København.
If I do end up typing lots of Danish, I'll probably map AltGr+A,E,O to Å, Æ, Ø. É is rare, so I'll still use the Compose key for that and German / Swedish names.
It converts latex to unicode where possible. It's pretty impressive how much of Latex can be replaced with Unicode today.
This program translates LaTeX markup to human-readable Unicode when possible.
Here's the default text from that webapp:
Basic math notations: ∵ A͡B + B͡C ≠ A͡C ∴ ∬∜x̅ ξᶿ⁺¹ - ⅜ ≤ Σ ζᵢ ∴ ∃x∀y x ∈ Â
Easily type in hundreds of other symbols and special characters: , ℵ, Œ, ⇊, etc.
Font styles support: 𝔹𝕝𝕒𝕔𝕜 𝔹𝕠𝕒𝕣𝕕 𝔹𝕠𝕝𝕕, 𝔉𝔯𝔞𝔨𝔱𝔲𝔯, 𝐁𝐨𝐥𝐝 𝐅𝐚𝐜𝐞, 𝓒𝓪𝓵𝓵𝓲𝓰𝓻𝓪𝓹𝓱𝓲𝓬, 𝐼𝑡𝑎𝑙𝑖𝑐, 𝙼𝚘𝚗𝚘𝚜𝚙𝚊𝚌𝚎.
Now type in this box and try it yourself. ⌣̈
For example (<^>! means the AltGr key):
<^>!4::Send „
<^>!5::Send “
<^>!2::Send ‚
<^>!3::Send ‘
<^>!+6::Send “
<^>!+7::Send ”
<^>!+8::Send ‘
<^>!+9::Send ’
<^>!-::Send –
<^>!.::Send …
Also to use Caps Lock as another Control key: Capslock::Ctrl
Or make windows stay on top of others, if even they lose focus: <^>!t::Winset, Alwaysontop, TOGGLE, Ahttps://www.x.org/releases/X11R7.7/doc/libX11/i18n/compose/e...
For example, em dash is Compose+minus+minus+minus, en dash is Compose+minus+minus+period.
It's also useful for foreign glyphs like Compose+a+a for the Nordic å, Compose+s+s for German ß (with Shift it becomes ẞ, which the new German orthography rules officially recognise!), Compose+quote+<vowel> for the various umlauts and Compose+i+period for the dotless ı (with Shift it becomes the dotted İ).
Just press Ctrl-Alt-Shift-U, type (part of) the name of the character and select from the list of results.
— (em dash), ︱ (presentation form for vertical em dash), ⤐ (rightwards two-headed triple-dash arrow), ﷽ (arabic ligature bismillah ar-rahman ar-raheem) are all easy to enter.
If you've ever seen an English-language page on a Japanese website that used weird quotes, this is probably why.
it starts with a regular parenthesis "(" but ends with a fullwidth parenthesis ")"(U+FF09)
is this a similar thing?
• It is semantically idiotic because it's an accent, not a character.
• It is visually annoying because you almost can't see the thing.
• It is bad for usability, because on non-US keyboards the accents are implemented as dead keys. Yes, accent + space gives you the character but that's really unintuitive for people who grew up expecting accents only over letters.
Sacrificing these bits of ASCII is fine by me, because the language is small enough, and I also allow Unicode. For example, curved quotes are allowed and can be nested or contain ASCII quotes without escaping:
// Character literals
‘'’
=
'\''
// Text literals
“Some "text" with “curved quotes”.”
=
"Some \"text\" with “curved quotes”."
For the sake of usability, of course, everything in the core language & standard library has an ASCII spelling, like in Perl6. I’d like for other languages to adopt this view as well. If new languages allow proper Unicode notation in some sensible places, then programming editors’ input methods will catch up, e.g., automatically replacing “->” with “→” or “\theta” with “θ” (like Emacs’ TeX input mode).Also, does anyone know of a reference for keyboard layouts from around the world that includes estimates of the number of people using them? I’ve tried to keep things relatively easy to type on all the major layouts I know of, but I don’t want to alienate anyone if I can help it.
By that metric, wouldn't & be too Anglo-centric, and # be too Euro-centric? There are layouts out there on which neither is readily available.
English is the lingua franca of programming, so it’s hard to avoid some Anglicisms (like ampersand meaning “and”, dot instead of comma for decimals, and English-language keywords) without going against strong precedents set by other languages. If I really wanted to be pedantic, I might use /\ and \/ for logical “and” and “or”—those spellings are the major reason that the backslash even exists in ASCII.
Edit: I forgot HTAB was actually part of ASCII. Oh well!
* The unaccented letter
* The backspace character
* The accent character
If this makes no sense to you, try to imagine a literal, physical typewriter. Windows line terminators also work with a similar principle.
(The typewriters I recall also didn't have a 0 or 1 key, you used uppercase O or I for these numbers.)
But I do suspect it had something to do with ascii compatibility, I don't recall what. Very little of unicode is accidental, there's usually some reason for whatever in it.
aa ah az ba bh bz ca cz ch
treating "ch" as a single letter that comes between c and d.
or
ab ah az b c .... z aa
treating aa the same as a separate letter at the end of the alphabet.
Then there's other rules for sorting that aren't directly alphabetic, like that names beginning with "Mc" should be treated as "Mac" or "St " as "Saint ".
"10 cats" should sort after "2 cats", not before it.
Anyone who tries to sort by just numeric ordering is doing it wrong.
http://www.typewriters101.com/uploads/1/7/6/6/17660651/s7662...
For acute or umlaut you could use a + backspace + ' or u + backspace + " (or the opposite order). For grave or circumflex, I don't think there was a solution. Write it in by hand?
Even worse, I remember one shop where the teletypes didn't have question marks, so people used capital P's instead.
HTML tags have nothing to do with less than or greater than, yet here we are.
In conclusion, it's simply convenient to use the standard keys we have and to use the symbols that are on it to mean something different from their original meaning in order to be able to express ourselves succinctly so that we don't have to spend so much time typing as we'd otherwise have to.
Of course you could always buy yourself an APL keyboard and write your programs in APL and use an APL REPL as your command line instead of using bash ;)
' PRIME (U+2032)
" DOUBLE PRIME aka inch mark (U+2033)
have their own codepoints
http://practicaltypography.com/foot-and-inch-marks.html
which describes implications for typesetting coordinates and other things:
118° 19′ 43.5″
118° 19’ 43.5” wrong (curly quotes, although it renders identical in some fonts)
118° 19' 43.5" right
But for anything related to contemporary typesetting on the web I recommend Practical Typography and especially the Type Composition chapter:
http://practicaltypography.com/type-composition.html#links
Including notes on quotes and apostrophes:
http://practicaltypography.com/straight-and-curly-quotes.htm...
An ʻokina (U+02BB, as found in "Hawaiʻi") is neither an apostrophe nor a left quotation mark.
Hey, imagine being able to nest strings without escaping! What a concept!
> For the constructs except here-docs, single characters are used as starting and ending delimiters. If the starting delimiter is an opening punctuation (that is (, [, {, or < ), the ending delimiter is the corresponding closing punctuation (that is ), ], }, or >). If the starting delimiter is an unpaired character like / or a closing punctuation, the ending delimiter is the same as the starting delimiter. Therefore a / terminates a qq// construct, while a ] terminates both qq[] and qq]] constructs.
And I guess you could count HTML in, too.
Nesting string literals without escaping is a somewhat poor concept, though. Firstly, what does that even mean? Given `abc `def' ghi', what is the string here? Is it abc def ghi or is it abc `def' ghi? Secondly, what if I want to just have an unbalanced ` character in the string data?
https://www.gnu.org/prep/standards/html_node/Quote-Character...
> Although GNU programs traditionally used 0x60 (‘`’) for opening and 0x27 (‘'’) for closing quotes, nowadays quotes ‘`like this'’ are typically rendered asymmetrically, so quoting ‘"like this"’ or ‘'like this'’ typically looks better.
Is this link saying I can quit using `QUOTES' in my Emacs-documentation? That style always struck me as odd :)
https://www.gnu.org/software/emacs/manual/html_node/elisp/Do...
We now have two characters for apostrophe and extra ambiguity for processing correct right single quotes. Great job not breaking historical documents Unicode.
Tell that to GCC:
/usr/lib/gcc/i686-linux-gnu/4.6/../../../i386-linux-gnu/crt1.o: In function `_start':
(.text+0x18): undefined reference to `main'
Looks good to me, by the way.> Where ``quoting like this'' comes from
I did it for a while out of a habit acquired from working with TeX. In TeX, it is the source code syntax for encoding quotes. Of course, it is lexically analyzed and converted to proper typesetting.
> If you can use only ASCII’s typewriter characters, then use the apostrophe character (0x27) as both the left and right quotation mark (as in 'quote').
It looks like shit in any font in which the apostrophe is a little nine, which is historically correct. What you want is a little "six" on one side and a "nine" on the other, or at least some approximation thereof. Even if the apostrophe is crappily rendered as a little vertical notch, it still pairs with a backwards-slanted `.
(The representation of apostrophe as a little vertical notch, I suspect, caters to literals in programming languages.)
> If you can use Unicode characters ...
then you should still stick to ASCII unless you have other good reasons to. ``Can'' is not the same thing as ``should'', let alone ``must''.
> For example, 0x60 and 0x27 look under Windows NT 4.0 with the TrueType font Lucida Console (size 14) like this:
The idea that people should change their behavior because of which font is default on the Windows cmd.exe console is laughable.
Why? Using non-ASCII Unicode characters acts like a nice canary for detecting character encoding issues. Besides, why would I purposely limit my text to ASCII? It doesn't even suffice for English, let alone almost any other language I use — including my native language Dutch, German, and Japanese.
> git commit and git commit-tree issues a warning if the commit log message given to it does not look like a valid UTF-8 string, unless you explicitly say your project uses a legacy encoding.
> git log, git show, git blame and friends look at the encoding header of a commit object, and try to re-code the log message into UTF-8 unless otherwise specified.
Why not? The firmware itself would usually have no reason to care about the details of a diagnostic message's encoding, whether that be ASCII or UTF-8 - it can mostly just treat strings as bags of bytes. There might be some byte values that are special (nul terminator, % for printf, etc.), but UTF-8 is a superset of ASCII and represents extended characters using only bytes with the highest bit set, so there will never be 'false positives' of the special byte values. Other than that, the bytes can stay uninterpreted as they go over whatever serial port or diagnostic protocol the device is using, until they eventually show up on - most likely - some sort of terminal application on a modern computer, which probably supports UTF-8 already. So in most cases it should 'just work'.
Of course, there are situations where it won't just work, such as if the firmware needs to display the diagnostic message on a screen (by itself), but from what I've seen those are the minority.
edit: As for Git, what's wrong with people writing log messages in their language of choice? (Other than the social issue of it making it harder for English speakers to use the codebase.)
> (The representation of apostrophe as a little vertical notch, I suspect, caters to literals in programming languages.)
"Historically", U+0027 has been used as all of an opening quote, a closing quote, an apostrophe, a prime symbol, an ʻokina, a modifier, etc.
So the historically correct thing is render it as a vertical notch so it looks non-horrible in all these uses, and render U+2018 and U+2019 as the "little nine" and "little six" symbols.
You don't need to speculate what the representation caters to; the Unicode spec actually does explain this (see Unicode 9.0 Chapter 6 Section 2)...
> The idea that people should change their behavior because of which font is default on the Windows cmd.exe console is laughable.
So your alternative is to change behavior because of which font is default on a system from 1984 which no one uses anymore?
All the apostrophes look like a little nine: in contractions like it's, and the possessive 's.
That's the character that was included in the American Standard Code for Information Interchange.
Image: http://www.worldpowersystems.com/J/codes/X3.4-1963/page5.JPG
The glyph appearing in the standard looks like a little 9. It is denoted as "APOS" in parentheses. A reference to it is made in A6.8, calling it "apostrophe".
Wikipedia's (https://en.wikipedia.org/wiki/Apostrophe) page refers to a vertical notch glyph as a "typewriter apostrophe". The normal non-typewriter apostrophe looks like a comma.
Okina? That indicates a glottal stop in some languages none of which are English, and so which were understandably not represented in the American Standard Code.
Yes, and that's the character that was immediately overloaded to used to mean a whole bunch of other things, because ASCII only included 95 printable characters, and did not include a prime symbol, an 'okina, or a left single quotation mark.
For that reason, U+0027 is not an apostrophe anymore. As the only ASCII character that can be used for a long list of uses, it's been massively overloaded, which is why Unicode currently defines U+0027 as a typewriter apostrophe and U+2019 as a real apostrophe.
Often it's the quotes which have been silently (automatically) converted to a visually similar (but functionally incompatible) character variant.
It looks terrible and to me it's a disgrace!
Example: it`s versus it's or it’s. (first one is wrong).
Benefits of typing and using typographer’s quotes directly in your JS/JSON/HTML/source:
1. No backslashes or other escape sequences needed!
2. WYSIWYG
3. Retina screens and gorgeous modern fonts mean that your sloppy quotes will look extra bad if you just use ASCII quotes
If you’re writing manpages, though, you should be using the -mdoc macros (https://manpages.bsd.lv/mdoc.html), which have “Dq” and “Sq” macros that wrap the arguments in double and single quotes respectively.
And of course I used ascii analogues to type these into HN :-(
But why, though? To the best of my knowledge, HN supports unicode quite well, including the following quotes: »«›‹„“‚‘ (available with the help of AltGr and sometimes shift from keys y, x, v, b when selecting the German keyboard layout on my computer).
I remember staring for a long time at the file when I first saw an m4 macro. My brain was telling, surely this has go to be a typo, but then everything worked as expected. Then I learned that's a proper way of quoting there.
Same probably goes for French and other languages with their own sets of quotation marks.
P.S. Major coincide I was googling this very question yesterday?