Semantic Linefeeds (2012)
rhodesmill.org
rhodesmill.org
My greatest annoyance with the style (which honestly is the sole reason I didn’t take it up a few years earlier) is bad handling of em dashes, of which I am fond. In normal English typesetting, there is a line-breaking opportunity before and after the dash, but no space; but in all markup languages that I know of that support soft breaks, the line break is equivalent to a space. Well, actually this is no longer normative for HTML (see https://www.w3.org/TR/css-text-3/#line-break-transform, which shows the problem for CJK, which doesn’t use spaces between words), but for now, all user agents still implement it as “line break → space”. Consequently, when writing in these markup languages, I can’t break the line after an em dash as I would like to, because it would change the output.
One of the features of the lightweight markup language I’ve been designing, then, is the ability to control soft break behaviour in order to correctly handle at least CJK and em dashes that were not otherwise paired with a space. (And I’m curious if anyone has similar cases not well-served by current rules.)
—⁂—
(If you’re not sure what I’m talking of: source:
Example: an em dash—
like this.
Expected result:> Example: an em dash—like this.
Actual result in the likes of HTML, Markdown and reStructuredText:
> Example: an em dash— like this.
(Another reason is that Swedish, which I also write, only has the en dash, so its nice that I can be consistent.)
The traditional syntax is particularly suited for newspapers who valued typographic density.
but wouldn't that still be the case if you were handwriting your newspaper and xeroxing it or something
where does the printing press come in
also though i don't think fine typography is mostly determined by newspapers
https://practicaltypography.com/hyphens-and-dashes.html
Handwriting allows one to vary the point (width) of individual letters. Printing presses do not afford that luxury.
what we nowadays call 'microtypography' and think ourselves very avant-garde for employing is ubiquitous in medieval illuminated manuscripts; every line is full of subtle variations in letterforms to better fit the available space
i don't remember ever seeing it in fiction
i think of space-en-dash-space as just being an error, and i'm pretty sure i wasn't just misidentifying en dashes with spaces around them as em dashes
Out of interest, why leading zeros?
What does your design for that feature look like?
For this particular aspect, I’m still undecided on how best to actually implement it. The most likely approach is to hard-code rules (possibly in a couple of groups, e.g. “CJK” and “other”), with the ability to opt into or out of them as part of dialect configuration (which is a bit like how you can define custom roles or change the default role in reStructuredText, but more general, able to change more aspects of the language’s syntax—things like change *…* to be something other than italic, make ~…~ strikethrough, define new counter styles). You could also generalise some form of declarative line-break-collapsing rule like “if preceded and succeeded (after whitespace trimming) by a character with Unicode property East_Asian_Width ∈ {Fullwidth, Halfwidth, Wide}” or “if the preceding line matches /(?!< )—$/” or “if preceded by U+200B (ZWSP)” (this last example borrowing from https://www.w3.org/TR/css-text-3/#line-breaking, basically applying the formatting rules in reverse for parsing, similar to what I’m proposing with the em dash and doing with Counter Styles; when laying out, ZWSP introduces a line break opportunity without adding a space, so when parsing ZWSP and a line break, you clearly shouldn’t add a space). But unless I can be convinced of actual value in generalising it, giving just one or two switches is likely to be the most sensible implementation, for complexity (this will probably be an optional feature, incidentally), manageability and performance.
2019: https://news.ycombinator.com/item?id=19256059 (11 comments)
2012: https://news.ycombinator.com/item?id=4642395 (35 comments)
It is annoying when you search for two words and they are not found, because they are on different lines
We just need better diff/merge tools that can handle text without line breaks. wdiff is installed everywhere for in-line diffs, but no one seems to maintain it. There have been patches sitting around for years: https://savannah.gnu.org/patch/?group=wdiff
Whenever I try to analyse my reasons for this, I find they're all bogus except perhaps one: with fewer linebreaks, there's a higher density of stuff on screen, which means I can refer back to more text as I'm writing. I don't think this is a particularly good reason in the age of marks, though.
______
† previous versions of this comment erroneously stated higher (and implausible) data rates
(one of the early hints that Donald Knuth was of the tribe which would become known as "geek" is that as a schoolchild he loved diagramming sentences)
The main advantage is that you get really nice diffs.
git diff —-word-diff