Ligatures in programming fonts: hell no (2019)
practicaltypography.com
practicaltypography.com
The article should have led with this. Instead this main caveat is buried in a paragraph, where it is preceded by, what boils down to, "do what you want, but you will recognize its wrong someday." Of course people will focus on that part and on the most prevalent use of ligatures in programming: In their own local editor. (Where people usually have strong opinions about what's the right thing to do.)
The tangent in (1) on how they contradict unicode could have been skipped as well, just go straight to the "the problem is" part. Why? Unicode already "contradicts" itself: △ is WHITE UP-POINTING TRIANGLE (trine) or U+25B3, which I am sure everyone can easily distinguish from the mentioned Δ Greek Capital Letter Delta U+0394 especially when used outside a clear context. (There are other symbols that are easier confused, but I find them close enough for the point, especially because fonts might render them differently.)
For private use, I find ligatures make spotting negations and reading comparisons easier more often than they are wrong. But I see why ligatures might not be a good choice for a presentation or code blocks on a website.
Not only because confusables already exist, but also because (as I said[1] the previous time this was posted) covering all ligatures used in all typographical styles is very much a non-goal of Unicode. The official position is that the font shaping layer[2] sits atop Unicode’s semantic representation and is free to ligate, spindle, or mutilate it for display however it prefers (at least for Latin, Greek, and Cyrillic it’s a preference; other scripts can’t be rendered at all without doing it, such as Arabic—barring the legacy presentational forms—or Burmese[3]).
The only reason Unicode even has those ligatures is that some IBM encodings (which were more presentational in nature) encoded them, and the IBM employees who wrote a large part of the early standard (based on the decades of i18n experience they had at that point) wanted roundtripping.
[1] https://news.ycombinator.com/item?id=29639966
(Alt Gr + m, at least in my language)
Unicode is a different code.
(and as I learned today, Julia code might be exclusive here)
If your point is that your program should use Unicode internally, this is already true for most programming languages but is independent of your input, plenty of languages routinely converting UTF-16 to UTF-8 to work with Linux or the web when they use UTF-16 internally, or UTF-8 to UTF-16 to work with Windows when they use UTF-8 internally. But I'd argue if you've done that work already, what harm is there in allowing your users to type € or á or 風?
Which was my point. The encoding might be different. For ASCII 7 bit is fine, for UTF-8 and only ASCII, 8 bits are required and, as you also already point out, the high bit must be set to zero with it.
So the point is to name the encoding (UTF-8) of the ASCII characters.
And the input was portable code, not user input of a program.
(Not making you look: U+0394 GREEK CAPITAL LETTER DELTA vs U+2206 INCREMENT)
I'm not sure if this is default behavior or something I did to it over the years.
I can see it being helpful but it feels way too trigger happy. It's always notifying me about Hebrew letters that only look similar to other things if you take off your glasses, squint, and change the font.
But, I was wrong. In several years of using them now, there has never been even one case of the theoretical "mistaking a different glyph that looks similar to the ligature" or "mistaking a ligature character sequence for a single character".
It's never happened to me, and I have never heard of it happening to anybody else that I know either. If you have a U+2260 in your code, and that's not valid in your programming language but your editor isn't flagging it as an error, then you have way bigger problems than ligatures.
On the flip side, though, it has absolutely made it easier for me to scan and read my own code, and the code of other people. So what I initially thought was "all downside, no upside" turned out to be quite the reverse.
Note the "for me" qualifier, though; I get that not everybody benefits in these ways. YMMV. I like speed metal, avocado toast, ligatures, and Allman brace style. You may not like these things.
But only that last one has any effect on others. If I insist on using my favorite brace style, sure that's annoying if you have to regularly work with my code and you prefer one of the incorrect brace styles like K&R.
But ligatures don't affect you; they just make my life better. Therefore, it's not a topic worthy of controversy, or of trying to tell other people whether or not they should be using them.
The main reason is I find the Unicode support much better. With Fira Code, some Unicode characters have varying widths even though it's supposed to be monospaced, making it sometimes look like a proportional font. JuliaMono doesn't seem to have this issue.
Example with Fira Code: https://i.imgur.com/YqlzgTE.png
Example with JuliaMono: https://i.imgur.com/41aHGVM.png
https://www.programmingfonts.org has a nice list, but besides JuliaMono only one other font on the list retains proper monospacing for less frequently used Unicode characters. You can test this by pasting the example below:
⋈ : naturaljoin
⋉ : leftsemijoin
⋊ : rightsemijoin
⟕ : rightouterjoin
⟖ : leftouterjoin
⟗ : fullouterjoinIf only that could be as easily configurable in IDE as liatures.
Rust’s language service (optionally) adds phantom code to several modern editors, showing inferred types in the editor, that aren’t really part of the source code.
That is very nearly the same thing...
I've been using programming fonts with ligatures for years. I had wasted several hours because of those stupid bugs when I mixed up O and 0, - and _, but ligatures never caused me any problem. It is obvious when they are there, and you rarely see non-ascii characters when programming (unless you are writing Lean).
Ideally, languages would support Unicode as well so you can mean what you say when you type—some folks will naturally prefer this to “keyboard-compatible” options and this should be fine (though ask many non-Anglophone users where the `$` key is and they won't be able to tell you). With Vim digraphs, Kitty unicode input, extra layers on a keyboard… there are many ways to tackle the ‘input hard’ problem. A good example of this is Dhall where Unicode is optional, but could be enforced with a built-in formatter option.
But a presentation introducing syntax? That's just silly. I can't believe someone would get that wrong.
This seems a much more reasonable take
Or if you're screensharing while pairing, or if a coworker is looking over your shoulder when you're sharing something, or...
$ nix-shell -p tmate --command tmate
Now they have it for a one-off $ nix shell nixpkgs#tmate --command tmateSorry, as a Lisp person, I don't see the problem. In Dybvig's typography, it's evident that an opening quote which looks like a 6 is the quasi-quoting backquote, and the closing quote is the regular quote.
It is legitimate for the apostrophe to appear as either a vertical tick or a 9 shape:
https://en.wikipedia.org/wiki/Apostrophe
Particularly this section:
https://en.wikipedia.org/wiki/Apostrophe#Typographic_form
Evidently, the 9 shape is the "typographic form" whereas the straight tick is "typewriter". The last paragraph of that section claims that this was a simplification for typewriters so that the glyph could fulfill multiple roles at the same time (including, possibly, simulating an exclamation mark, if overlaid over a period!)
In the USASCII system, there is only one character code for this, so in a given font, you get only one or the other. In a font where the ' character looks like a 9, you want ` to look like a 6. Or else if ' looks like a vertical stroke, then ` looks like a backward-slanted stroke of a size which harmonizes with that one.
It looks like the printed version of Dyvbig's book has a font which is typewriter-like, but uses the typographic form for these glyphs (thus undoing the hack that was perpetrated for the sake of typewriters). It's quite readable; there is no mistaking the forward tick for the backtick.
I now want a programming font which has those.
Minor issue mostly.
I tend to "no", but for no very good reason: Aesthetically they're lovely.
No, rather in your web page, in the slides that you're projecting, in the paper that you typeset into PDF...
The point is not about how you edit your own code. You could even write comments in sanskrit, if they are for your eyes only, but please don't dostribute them!
Hmm...my programming language supports ≠ directly ...
> even lower chance to somehow type it by hand.
... and my keyboard lets me type it (almost) directly, via Alt-=.
I think the author is spot on in that the current state is not the correct end-state, though given path-dependencies it is understandable how we got here.
If it were properly studied, and it turned out that ligatures have an effect on the rate of mistakes, then one side of the debate would have to concede that though they choose persist in their personal taste, it is in conflict with what is good for their work.
With all due respect, am I the only one expecting a strong argument for why ligatures on my own personal machine's terminal emulator should go?
I think "code for others to read" having to be free of ligatures was never in dispute at all. Unfortunately, the most significant reason isn't even mentioned: people's lack of awareness of the existence of programmer font ligatures. In clear terms: a confusion between != and ≠ does NOT exist for people who don't know of the existence of ligatures: they'll straight up believe they're looking at a ≠. (* if not slightly put off by the spacing around the character).
The one thing I did take away from this, is that people using Polacode/CodeSnap and the likes, should probably be aware to disable ligatures before exporting their pictures.
Ligatures in programming fonts: hell no (2019) - https://news.ycombinator.com/item?id=29637876 - Dec 2021 (195 comments)
Ligatures in Programming Fonts: Hell No - https://news.ycombinator.com/item?id=19805053 - May 2019 (86 comments)
Ligatures in programming fonts - https://news.ycombinator.com/item?id=14830797 - July 2017 (82 comments)
Luckily I rarely share my code through pdf, but mostly through git. Therefore, that hypothetical beast will read the code through its own editor, so the risk is mostly avoided. There is still the risk that this beast might appear behind my shoulders, see the pretty symbols and start that process anyway, but I have no idea how often that might happen too. It's a weird thing that this beast can do such an elaborate process, but not ask a question about what the symbols are.
For better arguments, there's a long history of problems with programming languages that use non-ASCII symbols as default, and it has been considered a mistake.
That crazy idea of 'letters' that look alike and thus introduce bugs is much worse with unicode than ligatures:
>>> a = "hello world"
>>> b = "hеllо wоrld"
>>> a == b
False
Example from "Unicode text spoofer":
>This browser-based utility fakes regular characters in text and replaces them with Unicode characters from other alphabets. The text that you paste or enter in the input text area automatically gets all possible characters replaced with Unicode homoglyphs in the output. You can spoof letters, punctuation marks, spaces, and even insert zero-width spaces between individual symbols.Now imagine instead of doing this, you use hello_world as variable name. You can thus interleave use of different variables that look alike, and have really buggy code.
Last argument, which is the one most important to me. Typing <= or != is ugly as shit. That is not how these symbols are supposed to look like, it's just a practical kludge. Ligatures are a simple solution to have both, that do not require a revolutionary way to change programming languages, and only engage yourself. To share print screens of code without ligature maybe configure a second editor with another font. That does not sound really hard...
For example, the text spoofer issue can be resolved in the text tools, like having variables with different names can be colored differently (this coloring scheme is better than having color just denote variable even without spoofing as it allows you to catch typos easily) Also you can have zero-width spaces highlighted
This, too, is entirely personal preference. I've tried it and found it resulted in too many colors for my brain to assign individual meanings to, and then I couldn't even benefit from syntax highlighting anymore because those colors also blended into the rainbow of text.
The differentiation benefit is not. Personal preference is what tool you choose to deal with spoofing, eg, you could also use a different font for different symbol ranges instead of colors.
> I've tried it and found it resulted in too many colors for my brain to assign individual meanings to
Well, you don't assign individual meaning to individual colors, but to color ranges which a given type can vary over. For example, going from green to yellow could be the range for variable names, so you associate that range with variables
Which color ranges have you tried?
I mean am I the only one that is browsing with completely disabled JS by default? :-D
More and more sites don't load anything withnout a JS and I'm not talking about advanced apps like Microsoft Office online...
But it's sadly true that more and more pages don't load anything if you have JS disabled. Often, the answer to "is this worth disabling JS for" is no.
Usually when I get empty page I just close the tab and move on. Very rarely the content is so unique and in general worth the hassle anyway. The thing is that you enable JS and it will load bunch of trackers and cookie popups and GDPR popups etc... I just hate that.
I wonder why you would self sabotage your website like that.
And I can only imagine, how screwed up internet must be for disabled people that need some kind of reader etc...
I have no patience for that. There’s enough frustration in my life without inflicting more on myself.
When I disable JS globally, I genuinely get faster experience on most of the pages.
The author has hidden their point all the way at the end of the article, though. They're mostly talking about sharing code, such as in screenshots, printed works, slides, etc.
People who don't know fonts commonly use ligatures (plenty of fonts that mess with fi or other commonly custom kerned characters) so they don't realise I'm not actually using the ≠ or ≥ characters in my code.
Unicode is an encoding tool, fonts are a display tool.
And just because the right symbol happens to be available in unicode doesn't mean it's right or easy to use it in every context ; especially when writing a compiler or when writing code. Who wants to have to write ≠ in their code dozens of time a day, without a keyboard made for this?
Fonts exist just for that: decide what the encoded text you have should look like. They are not competing with encoding, they're complementing it. Different fonts exist because people have different opinions about what's best for them to read.
The second point that is raised is a matter of tradeoffs and personal taste. If someone feels like seeing ≠ instead of /= or != in a given context is worth a surprise once a year, who am I to say they're wrong? It's not like I have to use their font, or like they can't use different fonts for different situations.
The rest are terrible. Especially if you don't have a math background. If you come from coding, the symbols are meaningless new inventions. To me, >= IS exactly the right symbol. It's the one I learned, the one I actually encounter in daily life, anything else looks like a math paper, not code.
Even emoji aren't the best if you didn't type them explicitly. :) is mostly similar to although not exactly the same, but :P is definitely nothing like any of .
Edit: apparently emoji don't work here again? I thought they added them?
Yes. On my computer. In my editor. That I am using…because I like ligatures, because they make my code easier to read. For me.
Use of ligatures on my setup does not force anyone else to use them? They’re not “viral” like a language feature.
“Oh but what if the >= Unicode symbol ends up in your source code and you don’t notice because you’re using ligatures?”
Is this the most contrived issue ever? Yeah and what if someone runs “find and replace” using Unicode confusables on my codebase? That’s arguably worse and requires no special font features. Plus I suspect most programming languages would throw an error immediately if you used the >= symbol instead of the pair.
If you don’t like ligatures, that’s fine, but there are those of us that do, and our font preference isn’t a threat to you.
Later in the submission, Butterick addresses this directly, and doesn't disagree:
> “What do you mean, it’s not a matter of taste? I like using ligatures when I code.” Great! In so many ways, I don’t care what you do in private. Although I predict you will eventually burn yourself on this hot mess, my main concern is typography that faces other human beings. So if you’re preparing your code for others to read—whether on screen or on paper—skip the ligatures.
> Bottom line: this isn’t a matter of taste. In programming code, every character in the file has a special semantic role to play. Therefore, any kind of “prettifying” that makes one character look like another—including ligatures—leads to a swamp of despair. If you don’t believe me, try it for 10 or 15 years.
The submission reads very much like an attack on any use of ligatures in code. The comment in the end reads, to me, more like a resignation of responsibility over his claims.
An argument, even one with strong language, can be about whether or not a tool is fit-for-purpose. This is allowed. This is discourse.
Right now, politically, a lot of arguments are actually attacks (on specific groups of people) in disguise, and I sympathize with your confusion. But arguments do exist. Not every argument is an attack.
I can't link to this book often enough: https://www.goodreads.com/book/show/29363252-conflict-is-not...
But metaphor is imprecision and unnecessary colour. Be clear and precise. Use specific language.
This is admittedly prescriptive usage, but I am prescribing this antidote for an overheated era, full of overwrought language.
But it is interesting still, how many seem to read the text it feels to them touching their privates and need to make themselves vent off.
Not dissimilar to those choosing to highlight the worst aspects of comments rather than addressing them in good faith. Please don't make things worse.
Here he's referring to code in books, slide decks, and during presentations, not for editing.
Just below the section the above quote is pulled from, Butterick explains
> One inspiration for this piece was the LaTeX crowd, who would routinely write me to insist their typography was infallible. And yet. I kept seeing LaTeX-prepared books that incorrectly substituted curly quotes for backticks.... “But code samples like these aren’t really ambiguous, because everyone knows that you don’t type the curly quotes.” ... But many of today’s programming languages (e.g., Racket) accept UTF-8 input.... So ambiguity is a real possibility. Same problem with ligatures.
Besides, if someone insists, it's also possible to make a language-specific setting in the editor to disable ligatures in those languages :)
All right, oh my captain, keep the indents and spacing tight on the monospace and sail along!
Btw., anyone still using monospaced fonts in their editors?
Tell me you've never programmed without telling me you've never programmed… Almost all code editors and IDEs in current use are set with a monospaced font by default, and the overwhelming majority of programmers keep it that way.
No judgement if that is good or bad or at all. Just a reminder that it ain't monospace any longer.
> Your code is still monospaced in all the ways that matter.
Just it ain't monospace any longer. (reminder)
In monospace one letter/symbol (printable character) has constant width.
The ligature is one letter/symbol. so while the stylistic decision retains the monospace style of the replaced symbols, the ligature itself is double (or triple) sized compared with other symbols.
IIRC this is commonly called "monospaced ligatures" ref https://github.com/tonsky/FiraCode/issues/712
E.g. a compacting ligature for "www" seems to be welcome in some typefaces for coding. But again, only to stress the reminder: It ain't monospace any longer.
We do still may prefer for the typesetting pleasantry that it still is styled as monospace (or fixed width). but those ligatures aren't fixed width any longer as with the letter A, that's all.
This is the difference between what it looks like (two or three glyphs) and what it is (a single glyph). So we break monospace to keep it appearing fixed-width even with the ligatures.
Would be interesting if there is more typesetting theory about the width of ligatures in a monospace font.
1 ASCII byte, 1 cell. That way ligatures don't change the width of the original text. If a line used to require 42 columns, it still requires 42 columns. That's why we call them "monospace": each ASCII byte requires the exact same width, even when it's part of a ligature.
> a compacting ligature for "www" —
Is obviously out of scope. We're discussing monospace ligatures here, not compacting ones.
The question is aren't you?
It's weird in a way, because there isn't all that much alignment going on in my code[1]. There's indentation, but that still works with a proportional font. And yet, it still looks wrong and hard to read. I can't really figure out why.
[1] I know some people align assignments in multiple lines, or comments at the end of the line, or similar things. I've been working under code style directives that prohibited those because lines changed more often because of whitespace than because of content. Since my work requires me not to do that, I don't have the habit at home either.
I did actually spend a year or two coding in a variable-width font, back in the '90s, but I switched back after the novelty wore off.
I'm not so sure if I want that, maybe I get dumb by it?!
My fav sites on this topic: https://pixelambacht.nl/2015/sans-bullshit-sans/ https://www.sansbullshitsans.com/
The former doesn't help in programming. The latter is fairly standard for ideographic languages, and has been well worked out for Chinese and Japanese.
Not really. Typing CJK text is a tedious and frustrating experience. People cope with it because no one has come up with a better alternative.
Both Chinese and Japanese alphabets have way too many characters to fit on a keyboard. To deal with this, CJK users type in the English spelling and convert it to CJK characters using software called input method editors (IMEs).
The problem is, this conversion process is often not straightforward. Unlike English, CJK sentences don't have spaces in them so IMEs have to guess how to split sentences into words. Then, it has to identify the right words, which is complicated due to the existence of homophones.
To deal with all the ambiguity involved in this conversion process, IMEs rely on predictions to provide users with a list of possible conversion candidates.
This leads to an awful typing experience in which users have to constantly choose from the IME-provided list for every little snippet of text. This is distracting, slows down typing, and makes key presses non-deterministic. To make matters worse, IMEs sometimes get things wrong and users have to get "creative" to work around it.
In comparison, typing English text is a breeze. You can just type what you want directly.
- A character harder to type
- The same potential confusion
- Code bases with both operators
Maybe programmers just need to know that ligatures exists. If you are programming in Java or Kotlin you know that intelliJ tend to use it.
Of course, if you have code, coping it to an ide with ligatures disabled will be rendered there old way just add it will use your preferred font or color highlighting schema.
So most of the problems are for images like slides.
Julia editor plugins let you do that - not an automatic replacement, but you can type the Latex for a symbol and press <Tab>, for eg. `\ne<Tab>` to get `≠`.
I suppose there could be an automatic-replacement mode, for those that have exact ASCII equivalents. So that, for eg., the editor replacing `!=` in your code with `≠` doesn't affect anyone else, because they're exactly equivalent and others can continue using `!=` if they prefer that.
U+2062 is the obvious solution.
"Huh? Why isn't this >= rendering as ligature?"
I tried for a while and it had pros and cons. Unfortunately it made me be confused one time too many so I reverted to not using ligatures.
Point is, find what works for you, we don't have to debate about this, and thankfully, we don't commit that stuff so that it becomes flame wars like tabs and spaces.
Would moving beyond ASCII be that much of a stretch for modern languages?
[1] https://docs.julialang.org/en/v1/manual/mathematical-operati...
• ≥: Compose > =
• ≠: Compose / =
You can define your own sequences too, if you want. I like to type exactly what I mean, including curly quotes and dashes and narrow no-break spaces and emoji and so on, but often didn’t like the default mappings (if there were any). As an example, I use this for curly quotes (look at the keys’ locations on the keyboard to understand—it’s coherent):
<Multi_key> <semicolon> <semicolon> : "‘"
<Multi_key> <apostrophe> <apostrophe> : "’"
<Multi_key> <colon> <colon> : "“"
<Multi_key> <quotedbl> <quotedbl> : "”"
(The only thing I can think of where I don’t normally type what I’d theoretically like to is hyphens, where I use ASCII’s HYPHEN-MINUS instead of the proper HYPHEN. I’ve never come up with a mapping I like—what with en dash, em dash, minus, &c.—and it’s also too common for fonts to lack a HYPHEN glyph.)When you’re used to this sort of thing, it comes completely naturally.
Somewhat more information: https://en.wikipedia.org/wiki/Compose_key
Not to mention there isnt one on your regular keyboard.
When you’re used to it, and given the comparative infrequency of the characters you use it for, it’s not much of a burden—just ~three key presses instead of one.
> Really dont need an extra modifier key.
Note that Compose isn’t a modifier key in the most common understanding of the word: you don’t hold it down while you press the others (which would certainly be awful for its ergonomics), but you press keys in sequence.
> Not to mention there isnt one on your regular keyboard.
On my current laptop I map RAlt to Compose. On my previous laptop, it was RMenu. There’s always at least one key there that you literally never need to use, and I find the ergonomics for thumb usage very comfortable. (But I really wish Space was split into at least two keys. It’s so stupidly large.)
If you can stand the tiny distant Return, get a Japanese keyboard.
I personally use AltGr (Mac Option) more than compose, since it's tractable to have almost the same layout on *nix and Mac. IIRC the default Mac layout has ≠ and ≤ ≥ in the obvious places, and they are just as easily touch-typed as shifted characters.
Compose is nice (on xkb systems) in that it's very easy to add personal customization to ~/.XCompose for task-specific characters.
For very common sequences that you always want to have mapped you can use a tool like AHK so that you don’t have to press Compose before.
The beauty of Compose though is that it works everywhere, and that it gives you access to a much larger character set in a very mnemonic way. For example, “ga” can map to “α” (“greek a”), “gb” to “β”, and so on. “2-” can map to “–“ (en dash), “3-“ to “—“ (em dash), “3.” to “…”, “c,” to “ç”, “i,” to “į”, “e,“ to “ę”, and so on. There are a lot of patterns that apply to a whole range of characters, it’s a bit like how commands combine in Vim.
help?> ≥
"≥" can be typed by \ge<tab>
search: ≥
>=(x, y)
≥(x,y)
Greater-than-or-equals comparison operator. Falls back to y <= x.
Examples
≡≡≡≡≡≡≡≡≡≡
julia> 'a' >= 'b'
false
julia> 7 ≥ 7 ≥ 3
trueA proper way would arguably be for the programming language to only understand “≥“, and for editors to provide support for entering such characters (or the user can configure a system-wide Compose key or similar). IMO we should work on standardizing keyboard input for those symbols instead of de-standardizing code display via ligatures.
Unclear if it will ever be ubiquitous, but I'm betting on widespread in a couple decades.
I've never actually thought about using comparative characters, but that actually makes a lot of sense. I think people will disagree if you use them, but there's no reason why you wouldn't be able to call a variable x≥y≤z…
And the most basic reason is because a dedicated symbol looks much better (and doesn't break vertical alignment since it's one glyph)
You could even compile a bunch of them into a header file.
life ← {⊃1 ⍵ ∨.∧ 3 4 = +/ +⌿ ¯1 0 1 ∘.⊖ ¯1 0 1 ⌽¨ ⊂⍵}
All this debate about individuals and their text editor setups is inconsequential compared to that.
It’s like a rectangular sink with sharp edges that’s hard to clean. It’s like writing javascript without semicolons: yes, it’s prettier, but you gotta keep yet another set of rules and exceptions in mind to get it right.
For example, a JS regex like /=foo/ combines the /= into a ligature in Fira Code with default settings. That’s just confusing right? The / is a delimiter and the = is part of the regex itself. It means you need to keep like 0.5% of brainspace available for “reverse parsing” the ligatures on your head. Why waste that? Why spend even a tiny fraction of your mind worrying if you saw that thing right?
Ligatures which are ligatures and thus show that the unit is a semantic unit, but displays a minimal visual break in-between, showing that there are two code points underneath. Similar to the classic fi ligature which doesn’t make a complete new glyph, but arranges the glyphs for f and i in a more pleasing manner.
For programming ligatures that would mean that "===" doesn’t give "≡" but more something like "⩶". I’m approximating with mathematical unicode characters here, the real thing would of course have a width of 2 or 3ch.
In JuliaMono/Fira Code, if you sequence `::`, the left colon in shifted slightly to the right and the right colon is shifted slightly to the left and both are scaled down slightly vertically.
Try ::, :=, --, #!, ``, !!, ;; using Fira Code: https://www.programmingfonts.org/#firacode
These fonts can contain optional features which can be enabled/disabled dynamically by supporting editors. You could customize the behavior for specific languages and have the editor only show those features when you have the relevant language file loaded.
Ideally I'd like to be able to e.g. highlight one of these weird ligatures, or a whitespace that I want to identify, or an emoji my font doesn't support and select something like "identity this character" from the context menu to see what symbol or series of symbols I'm actually looking at.
If you're OK with a command line tool, I just finished updating and finally publishing a Python script I wrote for that purpose around 15 years ago: https://pypi.org/project/unicode-name-inquiry/ Lots of options, but default usage is, for example:
$ uni ≥
≥ U+2265 GREATER-THAN OR EQUAL TO
(It doesn't currently do the string-of-symbols case, because it tries to interpret strings as some other character identifier, e.g. `uni tilde` or `uni 8805`, but that sounds like a reasonable option to add.)Typically, people just search for the character on the web and land on one of the many ‘monetized code charts’.
I've been teaching myself python on the side in order to automate parts of my job so it would be neat to see it used here.
In text editors you can only go all in or not. It is not really intended for that.
I don't think that needs to be the case. Text editors already do syntax highlighting, and if using treesitter they have accurate understanding of the syntax. So there isn't major blockers for applying ligatures selectively based on syntax.
Other than that, ligatures are great and I will never turn them off.
That's just silly, the reason there are ligatures is precisely because all these programming languages are poorly designed not to allow using the proper ≠ instead !=, but then this is also why there is no ambiguity - your editor would highlight the real ≠ as a syntax error
Why? This brings us to refute of point two: lack of semantic substitution is a hypothetical problem (it could easily be solved though by parsing the source before substituting).
But here is the refute: how often do I come across this?
And how, exactly, could I shoot myself in the foot with it?
The author omits answering that question.
I studied typography alongside CS and worked in the former field for years so I have strong opinions on use of type for stuff other people look at.
My editor is, however, something only I look at.
I use code-type fonts with ligatures since years in various editors (vscode the last four), writing mainly Rust. I never ran into either issue one or two.
How so? It's either a ligature or a syntax error: weird unicode symbols can't be valid operators and operator ligatures cannot be valid identifiers. The exception would be string literals, though that situation would be quite rare - probably less common than ambiguity due to different but similar-looking unicode characters. Though I agree it mostly makes sense in one's own editor, not when sharing the code.
Meanwhile people don't mind at all that their code may contain invisible characters with variable length that can make cursor movement less predictable and cause syntax errors - tabs.
As for "people don't mind [tabs] at all", some people mind very much and consider them pollution. If you mean _some_ people don't mind tabs, ie you're highlighting that tabs vs spaces is a bigger deal than ligatures, I concur.
The logical reason to prefer spaces is that the choice is not "spaces vs tabs", it's "spaces vs (tabs and spaces)", and that mix of (tabs and spaces) is highly problematic.
Yes, contrary to spaces/tabs ligatures only affect how it is displayed on screen, letting everyone choose how they want it without any drawback or inconsistency. Which is why I find the other argument in the link a bit unconvincing too - most code isn't shared in static screenshots, though I'll agree when it comes to books.
1) the last sentence of the article implies that the author of the article abhor them as much as programming ligatures. I don't understand why but preference in taste, color, esthetic are not objective, nor absolute, so I am not the one to judge him.
I am not very familiar with the latest trends, could anyone else chime in here? Do Julia coders use fonts with ligatures? That could indeed be one case where confusion is possible. Or maybe not?
I personally don't like any ligatures, especially when people use them in presentations to public like the article mentions. Confusion between operators is rare, but possible, for eg. when using Catalyst (https://docs.sciml.ai/Catalyst/stable/catalyst_functionality...).
> It seemed to me that the use of Greek letters and symbols is actively encouraged in the Julia documentation.
I don't remember getting that sense from the docs, but I may be wrong. The general convention is that Unicode at the level of user code is fine, and a lot of the community (including me) likes it because it makes it much closer to the actual equations and scientific notation we're working with. For libraries, it's strongly encouraged to have ASCII equivalents for any Unicode operators, keywords, etc. that you expose, and all the major libraries do this. The same is true in base Julia as well, every non-ASCII Unicode name has an ASCII equivalent too.
This is in contrast to ligatures: say if there's a fi ligature, I type "<" "=", and the "≤" character is substituted during rendering, but the individual characters are still the "character string". (You could also choose to directly insert the "≤" character directly, and that would then be a single code point, but I think for the purposes of this discussion, people are mostly talking about cases such as when the character string "<=" is rendered as "≤".
If Julia (like APL) uses characters outside the ASCII range as syntax, I suspect that there are input methods to assist in entering them as opposed to ligatures representing them on screen. But that's just naïve supposition on my part.
Another interesting case is TLA+ [1], where the language shares some LaTeX operators which can be rendered for display.
What?
> The problem? Many of the programming ligatures shown above are easily confused with existing Unicode symbols.
Oh. But that’s not a contradiction? Ligatures are just for display and don’t really interfere with Unicode.
I don’t see how you could be confused, really, considering the state of the art. Widespread Unicode in programming languages are rare. You probably know if you are in the process of reading Haskell code (ASCII) or Agda (Unicode).
Undoubtedly he made an argument for ligatures, but why are Powerline characters in the same basket?
> The other inspiration for this piece were the people who repeatedly asked me when Triplicate would get ligatures, Powerline characters, and so on. Answer, as nicely as possible: never.
So the answer is, for the same argument. Not going to happen with that typeface, by the authors design.
(2) context is super easy to apply, an editor can very reasonably enable or disable ligatures depending on where they are in the code: sublime text does this already
If you’re printing code for others to read, stick to standard characters. If you’re choosing a font for your editor, do what you want.
I really like ligature fonts. I was an early adopter of Fira Code and later moved to JetBrains Mono. I like it because I find it’s much easier to skim and understand the intent vs a standard character representation.
At the end of the day, it’s your editor, do what you want.
So it goes to show that waiting for things is, like, worth it.
But there's a lot of bad things in this world, dude.
Like fucking skunks, dude. HELL NO.
Scratching your eye, but it's still fucking itchy. HELL NO.
The fucking Cubs, dude. HELL NO!
Fucking ligatures in fucking programming fonts, dude. HELL! NO!
BUT BANANA BREAD, ON FUCKING WORK, DUDE! HELL YEAH! HELL YEAH, BRO!