HN mining or: How I learned that there are some among us with soft quotes
github.com
github.com
Unfortunately, users (and product managers of competitors) misinterpreted the label in the check-box in the UI that enabled this feature, and thought that "Smart Quotes" meant that the name of the quote characters somehow had the word "smart" in it. But, no, it just meant "turn on the intelligent logic for which quote to insert into the document."
Source: Me, 1988 or so, working on FrameMaker. I don't think there's a reference to the word pair "Smart Quotes" anywhere earlier than the FrameMaker ~2 manual, and certainly not in any book on typography. I'd be happy to learn that I am wrong, if anybody has a pointer.
(FrameMaker was great, BTW. It was what got me into developing fancy layout style sheets and templates, which I later also did with Interleaf, MS Word, LaTeX, and various Web stuff. Somewhere in there, I worked on DynaText and DynaWeb at EBT, so overlapped a bit with Frame as it got into SGML.)
It's a shame that ASCII (and its visual representation on terminals from most manufacturers) has a wonderfully slanty left-quote character, but the right-quote character is typically shown as vertical. This makes TeX input files that contain quotations look kind of yucky, which is all the more galling as the original SAIL terminals had symmetric quotes, so everything looked beautiful there (never mind also having proper uparrow and downarrow characters to indicate superscript and subscript, rather than the ugly caret and underscore characters that we're all stuck with in ASCII-land TeX).
Source: Me, 1980 or so, working on TeX and Metafont.
And of course there’s the overload of hyphen and minus, such that Unicode gave up and named U+002D “hyphen‐minus,” and created two separate characters U+2212 “minus” (−) and U+2010 “hyphen” (‐) to use in situations where the typography matters. Most fonts seem to use identical glyphs for hyphen‐minus and minus, but I’ve seen some that try to split the difference, giving hyphen‐minus a glyph with a height and width somewhere in between hyphen and minus.
If you find minus and hyphen‐minus hard to distinguish visually, just remember that a true minus sign looks just like a plus without the vertical part. Compare the alignments:
- example (U+002D hyphen‐minus)
+ example (plus)
− example (U+2212 minus)
Those HTML entities are then curled correctly when any theme[2] is applied.
What bytes is not having a separate glyph for a curled apostrophe because it makes detection of British (or nested) quotations a difficult chore for natural language processors. To get a curled apostrophe, word processors inject the semantically incorrect right-curled single quote, which ought to be reserved for a closing quotation mark exclusively.
There are a few fonts that do curl the apostrophe, such as GFS Didot[3].
[0]: https://github.com/DaveJarvis/keenwrite
[1]: https://whitemagicsoftware.com/keenquotes/
[2]: https://github.com/DaveJarvis/keenwrite-themes/blob/main/xht...
See Markus Kahn's page at https://www.cl.cam.ac.uk/~mgk25/ucs/quotes.html which explains it well. (Also: the one-page documentation (plus one page implementation) of the LaTeX “upquote” package: http://mirrors.ctan.org/macros/latex/contrib/upquote/upquote... )
Of course, your links are right: it makes no sense to overload these characters anymore now that we have character sets and software capable of providing typographic niceties. But it’s fun to know the history.
For starters, most easily at hand, see Communications of the ACM, Vol. 8, No. 4, April 1965, page 207, which shows them by those names.
Of course, there are details. It says that the 0x40 character can also be used as a grave accent, but only if "preceded by an alphabetic character and a BS (Backspace) in that sequence." Oh, for the days of hard-copy terminals.
And 0x27 is worse. It is documented there to be used as Apostrophe, closing single quotation mark, and even acute accent (but again, only in the context of the three character sequence alpha-backspace-accent).
Fine, that was 1965, and no controversy. But following the actual ANSI standard along over the years, by 2007 they're still there in "Coded Character Sets - 7-Bit American National Standard Code for Information Interchange (7-Bit ASCII)" ANSI INCITS 4-1986 (R2007) / ANSI X3.4-1986 (R2007), Approved June 14, 2007; Table 7, page 12. (Well, in 1986 they'd been slightly renamed to be Left and Right Single Quotation Mark, so the names would make sense for right-to-left languages, too; just as Paren, Bracket, and Brace pairs had all also changed names from "Open/Close" to "Left/Right".) The subsequent revisions to the ASCII standard, of June 15, 2012 and November 2, 2017, I don't have copies of, but these character designations clearly lasted at least into the "updated definitions" current as of mid 2012.
So, left and right single quotes were clearly in the official standard for almost five full decades beyond "early ASCII". I expect that they survive still, and will soon go into their seventh decade, but happy to hear otherwise from anyone with access to the later standard documents.
(And, yes, of course the Unicode standard has its own (no doubt better) opinions about what goes on with the first 128 characters out of tens of thousands, and what they're called and what they're properly used for; but that's a different issue than what the ASCII standard itself has to say about cramming stuff into its overloaded character set with a grand total of 128 slots.)
A better way of stating it is to start with appearance: people who started with computers where 0x40 ` and 0x27 ' looked symmetric (like ‘ and ’) would likely end up using them as symmetric quotation marks (thus the conventions used in some software like Emacs and TeX), while everyone else, where they were not symmetric (looking slanted and vertical respectively) would likely not be led down that route. It's a bit unfortunate that there was this divergence, and that it wasn't settled early enough.
An interesting example is The TeXbook itself: the 1979 TeX manual (in the "TeX and METAFONT" book), the precursor to The TeXbook and apparently written with mostly the SAIL environment in mind, simply says (early in Chapter 2):
> In the first place, there are two kinds of quotation marks in books, but only one kind on the typewriter. Even on your computer terminal, which has more characters than an ordinary typewriter, you probably have only a non-oriented double quote mark (") because the standard “ascii” code for computers was not invented with book publishing in mind. However, your terminal probably does have two flavors of single-quote marks, namely ’ and ’, which you can get by typing ` and ´. The second of these is useful also as an apostrophe.
While already in 1984, The TeXbook (second printing, October 1984: the earliest I have access to), written with more awareness of the rest of the world, has more-or-less the same text, but adds:
> American keyboards usually contain a left-quote character that shows up as something like `, and an apostrophe or right-quote that looks like ' or ´.
So already the appearance did differ across different systems.
There's a bit more on another related page of Kuhn at https://www.cl.cam.ac.uk/~mgk25/ucs/apostrophe.html (this one is about the opposite problem: Europeans using acute accent as apostrophe), and there's a bit more about history of these two ASCII positions at https://jkorpela.fi/latin1/ascii-hist.html
If you’re going to be pedantic, I’m not sure that “left-quote” or “right-quote” is correct either. Shouldn’t it be “left double quotation mark” or “right double quotation mark”?
But see also https://smartquotesforsmartpeople.com/ :-)
“一個句子”
Left: “
Right: Shift+“
Typewriter: AltGr+“In that sense I think "straight quotes" are fine too, although “smart quotes” arguably look a bit nicer, but that's a matter of personal taste IMO.
Related comment from a few months back: https://news.ycombinator.com/item?id=31792626
To all those who are confused, smart quotes or soft quotes are those context aware quote marks that some word processors insert which depending on context, will put either closing or opening quote marks.
On PC I have AutoHotKey bindings for German and English quotes, single and double, and I'm usually using them.
Most linux desktops are different and more complicated but here is how I set mine up on openbsd.
pick a key that is handy but you don't use much, for me this was the context menu key, on many european keyboards there is an altgr key set aside for this exact use case.
in ~/.xsession add a line like "xmodmap -e 'keysym Menu = Multi_key'"
create a file ~/.XCompose to setup your own compose sequences. mine looks like
#load system compose files
include "%L"
<Multi_key> <w> <e> <b> : "\xf0\x9f\x95\xb8" # spiderweb
read /usr/X11R6/share/X11/locale/en_US.UTF-8/Compose to get a feel for the existing compose sequences.According to an ancient convention.
Conventions can be changed. It is not intrinsically, objectively superior to have two separate glyphs for a starting and ending quotation mark. I'd argue the opposite. It is needless complexity. Let us jettison it and embrace the simpler single glyph.
Mobile devices appear to do this by default. iOS in my case anyway. I turned them off because it was messing up my hugo blog's frontmatter whenever I start a post on mobile.
“ = option + [
” = option + shift + [
‘ = option + ]
’ = option + shift + ]
There’s a few others I use regularly as well:
… = option + ;
– (en dash) = option + -
— (em dash) = option + shift + -
Then there’s a bunch of accent meta keys, so you can use option + c for façade, option + e, e for fiancé, and option + u, i for naïve.
And specifically, "nouns" here are defined by this regex: https://github.com/chapmanjacobd/library/blob/main/scripts/m...
iOS at the least by default will substitute "dumb" quotes with smart quotes in text boxes.
This is a double quote: ", 0x22, ascii
This is a smart open quote: “, 0x201c, unicode
This is a smart close quote: ”, 0x201d, unicode
By default your keyboard double-quote key produces the ascii double quote mark, and that is what's used for quoting in virtually every programming language.
Word and many email clients will use smart-quotes by default, but they're pretty rare in programming, hence why the "some among us" phrasing calling them out as strange.
That's what the linked repo is referring to by soft quotes, and the command given specifically uses 0x201d above.
And I don’t see how one is any “softer” than the other.
People’s programming editors often do paren matching - does that make them smart parens? No they’re still just regular old parens.
Language is descriptive, and people do in fact call them smart quotes. That doesn't mean there has to be anything smart about them, and I'm not making the claim there is, I'm just using an agreed-upon term to convey what I'm talking about.
Cause it's shaped like a simple boat. Or sort of is. Kind of like how 'soft quotes' are considered soft because they aren't in a very hard like up and down position, and are instead positioned like a feather softly falling. (or so I guess)
Because historically, the insinuation was that cheap sausages contained dog meat, that’s why. Someone has answered your gravyboat question. And your ladybug example is bordering on blatant troll baiting.
> Language is descriptive, and people do in fact call them smart quotes. That doesn't mean there has to
Some “people” call them ATM machines even though the M stands for machine.
> I'm just using an agreed-upon term to convey what I'm talking about.
The agreed upon terms are Left Double Quotation Mark and Right Double Quotation Mark - nothing about smart quotes in the Unicode tables.
I get that the term smart is used colloquially - but your comment suggests that their Unicode names are “smart”. I do think it is overall silly and confusing to refer to the characters themselves based on some formatting or automatic insertion feature.
i.e. - the quotes are smart because they adapt to your region.
‘|’|“|”
but there are also other quoting glyphs in ASCII but especially UTF-8For example: 10° 33' 19" -> 10° 33« 19». Very smart!
A PM somewhere probably thought "smart quotes" sounded neat, and here we are.
[a]: On my computers, I’ll use "dumb" quotes. Easy way to see where I typed a comment from, I guess.
If someone can explain why it's not just egotism, please go right ahead. Otherwise, it really does seem like it. To me at least.
Other comments try to sort of deal with this problem in their explanations about soft and hard quotes, but they don't really deal with the part about it being disturbing.
What's so disturbing about using a feature available on your device? Especially when it does nothing harmful to the user, or recipients of that users communications?
Further, if you check my sibling comment you'll see that it is a common mistake for beginner/intermediate coders to paste these quotes into code. This is the source of the "cultural debate" in this instance. They are making a judgement about how this common copy/paste problem shows up in practice. Not a judgement about "proper usage".
It's honestly a little surprising that you took this so personally. It reads rather playfully in my opinion.
I never knew you could do that! I've got very frustrated sometimes typing on an iPad when I really want a dumb quote, and, unlike on macOS where the dumb quote is the default and there's the convenient if baffling OPT-[ shortcut to get “ and OPT-SHIFT-[ for ”, I couldn't figure out what to do to my default smart quote to get a dumb quote. Thanks!
Yes, I think it is a thing where people type into Microsoft Word (or similar). Given how moving around the web can cause text in fields to be lost, I don't blame them. But it was a surprising discovery for sure.
This is simply not true. Look at the files in the repo directly
Intel is ^Intel$. The nouns are separated by word boundaries
'sun|moon'
resulted in 30905 moon
36100 Samsung
44937 sun
my guess is no.