When monospace fonts aren't: The Unicode character width nightmare
denisbider.blogspot.com
denisbider.blogspot.com
JOE uses 4-level radix trees for character classes. These work well because the leaf nodes are highly redundant and can be merged together. The resulting structure is often smaller than a binary tree. Character classes are also used for regular expressions, so there is code to build them on the fly from a list of ranges (it's tricky to do this efficiently).
Anyway, I'm surprised that emoji are not double-wide characters.
JOE is still missing Unicode normalization for string searches.
When you're trying to align lines of a block, you only need that a string of n initial spaces (or tabs) is the same width every time: you don't care if it is the same as a string of n arbitrary characters. This suffices for the compiler-enforced indentation rule in Python (though not, I think, for the indentation rule in Haskell).
(I code in a variable-width font in Emacs and I work on shared codebases. The codebase style does have a couple alignment rules that don't make sense for variable-width fonts, but I just let Emacs enforce them and I'm happy that overall the code as I see it is easier on the eyes than it would be with fixed-width formatting.)
Visual Studio defaults to Consolas, which is nice.
Separately from the "lining things up" argument, though, there's an argument to be made for characters in any editing situation to be wider, even if they're not all the same width. For naturally-narrow letters like I and l, and for periods and commas, and of course spaces, the non-monospace fonts often make them so narrow that they become harder to target with a mouse, harder to distinguish from each other, and in some cases hard to even see at a glance whether the insertion point is to their left or right. This is a UI problem with a lot of editors (e.g. text entry for data or comments on the web) that doesn't get enough attention; if the content and presentation can be separate (i.e. it's not a WYSI-more-or-less-WYG editor) then the editor should use a font that doesn't have super-narrow anything. Monospace is a convenient way to achieve that.
For example, instead of this:
someObject.someMethod(oneArgument,
anotherArgument,
oneMoreArgumentForTheRoad);
Format it exactly as you would if the parentheses were curly braces: someObject.someMethod(
oneArgument,
anotherArgument,
oneMoreArgumentForTheRoad
);
Or instead of this: string someVariable = oneThing +
anotherThing +
yetAnotherThing;
Do this: string someVariable =
oneThing +
anotherThing +
yetAnotherThing;
The code is just as readable this way, and you gain many benefits:* Your code looks exactly the same in a proportional or monospaced font, so people viewing the code are free to choose either.
* You no longer need to fiddle with adding or removing spaces when you change the length of one of your variables or function names.
* The diffs in your revision history no longer show spurious changes that result from that column fiddling.
* Your code lines become much shorter.
* Instead of having one formatting rule for curly braces and a completely different rule for parentheses, you use the same formatting rule for everything.
* It doesn't matter if the code is indented with spaces or tabs. With no column alignment, rules like "tabs for indentation, spaces for alignment" are no longer needed.
* If you need to change the code to a new indentation style (e.g. change 2 or 4 spaces to tabs, or vice versa), you can do that reliably with a simple regex search and replace, with no damage to the code formatting and no manual cleanup required.
I adopted this practice long before I switched to coding in proportional fonts - and once I started formatting code this way I realized that it didn't matter any more what kind of font I used. I was free to choose any font that pleased my eyes, with no impact on the readability of the code for those who prefer monospaced fonts.
A good place to see the problems that column alignment causes is the Servo source code, whose coding standard mandates column alignment. I posted a few examples here:
https://news.ycombinator.com/item?id=9469713
I've often wondered why column alignment is so popular given its drawbacks. I have a theory that I think explains some of it, at least in the case of function calls and parenthesized expressions.
It's because of the very common objection to putting spaces inside the parentheses. For example, PEP8 and many coding standards explicitly forbid this.
Here's what happens. Take my first example above, written as one line without spaces inside the parentheses:
someObject.someMethod(oneArgument, anotherArgument, oneMoreArgumentForTheRoad);
That line is too long, so let's fix it. The natural starting place is to change every space to a newline: someObject.someMethod(oneArgument,
anotherArgument,
oneMoreArgumentForTheRoad);
That's a bit messed up, so what can we do? Indent the extra lines? someObject.someMethod(oneArgument,
anotherArgument,
oneMoreArgumentForTheRoad);
Ugh. Now the arguments don't line up at all. It's really ugly. The only cure is to align the columns: someObject.someMethod(oneArgument,
anotherArgument,
oneMoreArgumentForTheRoad);
But what if we adopted the practice of putting spaces inside the parentheses? someObject.someMethod( oneArgument, anotherArgument, oneMoreArgumentForTheRoad );
If we make the same substitution, changing each space to a newline, we get this: someObject.someMethod(
oneArgument,
anotherArgument,
oneMoreArgumentForTheRoad
);
And now it makes perfect sense to indent the arguments: someObject.someMethod(
oneArgument,
anotherArgument,
oneMoreArgumentForTheRoad
);
This is the same thing we intuitively do with curly braces. After all, hardly anybody codes like this: while(true) {oneStatement;
anotherStatement;
oneMoreStatement;}
I think one reason is that we pretty much tend to put spaces inside the curly braces, even in a one-liner. It's not too common to write this: while(true) {oneStatement; anotherStatement; oneMoreStatement;}
Instead, this seems more typical (if you use a one-liner at all, which I'm not arguing for or against): while(true) { oneStatement; anotherStatement; oneMoreStatement; }
No one seems to mind spaces inside the braces, but I've had developers totally freak out over the idea of putting spaces inside the parentheses. It Simply Is Not Done.I don't understand why there's such an objection to that, especially when it seems to lead directly to the difficulties of column alignment.
I tend to think that it's because spaces never go inside the parentheses in English text. But this is code, not English, and we are free to choose conventions that benefit us, even if they differ from how we'd write prose.
Edit: I added a number of thoughts after first posting this. If you upvoted this an earlier version of this comment and now think I'm insane for advocating spaces inside the parentheses, let me know and I'll post another comment that you can downvote. ;-)
if ( $test[0] ) func1()
elseif ( $test[1] ) func2()
elseif ( $test[2] ) func3()For the example you posted, I would classify that as column alignment rather than indentation. Just to clarify my terminology, what I call indentation is something that happens at the beginning of a line only. Any additional spacing after the first nonblank character is column alignment.
So the spaces before the if and elseif are indentation, and the extra spaces within the lines are column alignment (in my nomenclature).
The column alignment in your example is pretty appealing, but it does have some of the same problems as other forms of column alignment. When we get to ten or eleven tests, things have to get juggled around again. Do you do this:
if ( $test[0] ) func1()
elseif ( $test[1] ) func2()
...
elseif ( $test[10] ) func3()
or this? if ( $test[0] ) func1()
elseif ( $test[1] ) func2()
...
elseif ( $test[10] ) func3()
or go all-out with something like this? if ( $test[ 0] ) func1()
elseif ( $test[ 1] ) func2()
...
elseif ( $test[10] ) func3()
I think basically I am just lazy and found all of this alignment so tedious that I looked for ways to avoid it. :-) module fred(
a,
b,
c
);
...
endmodule
I've not tried it for function calls, but it's not a bad idea...How does it look for the case of an indented block after a while or a for? Let's see:
while (
a &&
b && c
) {
a = a + 1;
b = b - 1;
}
Hmm, maybe better than what I do now... while (a &&
b && c) {
a = a + 1;
b = b - 1;
} while(
a &&
b && c
) {
a = a + 1;
b = b - 1;
}
That's not a big difference; the only reason I leave out that one space goes back to my idea of having fewer formatting rules and special cases. I don't put a space before the open paren on a function call, so I don't put one there on a statement that has parens either. Not a big deal either way, I'm just lazy and like to have one rule instead of two. :-)I experimented in the past with some other ways of formatting this kind of code, e.g.
while (
a &&
b && c ) {
a = a + 1;
b = b - 1;
}
Of course you can see the problem here - there's no visual separator between the while expression and the statements in the loop body.That's why some coding standards require a double-indent:
while (
a &&
b && c ) {
a = a + 1;
b = b - 1;
}
That never appealed to me at all, but when I started moving the close paren to the next line and dedenting it there was no need for the double indent.What about this?
int c[] = {
3, 6, 9,
12, 15, 18,
...
300, 303, 306
};
or int v =
((r << 24) & 0xff000000) |
((g << 16) & 0xff0000) |
((b << 8) & 0xff00) |
( a & 0xff) ;It's very different from the examples of column alignment I was critiquing in my earlier message, where there's an equally good or better way to format the code using indentation only.
All in all, I still prefer the visual appearance and compact nature of proportional fonts, especially my current favorite, Trebuchet MS - and they do just fine for almost all the code I write. But for all my criticism of column alignment, I have to admit I do miss it in cases like your second example where it brings clarity.
My dream editor would have a way to allow the use of both proportional fonts where they work well and monospaced fonts where they are better, all within the same file. Perhaps a special commenting convention that would let you switch back and forth.
One of the editors I use, Komodo, has a bit of this. When you create a custom visual theme, you can specify both a proportional font and a monospaced font, and either toggle the entire file back and forth with a hotkey or select one font or the other as part of syntax highlighting - similar to the way many editors let you select bold or italic for particular syntax elements. For example you can select a monospaced font for comments so you can do ASCII art there, while using a proportional font for other code.
Thanks for sharing those examples!
You might be interested in literate programming:
https://en.wikipedia.org/wiki/Literate_programming#Tools
It's not exactly what you are after, but I think is more fundamentally what you are after - ways to make it easier to understand what code is doing. You can have the best of both worlds too here with something like noweb and two windows.
One other thing I ran into is a code base that had incredibly long lines (more than 132 often). So I started using ndiff to help me there. It's a python script that puts + and - under the lines where the differences are as well. I also wrote a little script to html it with different colors. It was pretty awful until I made it monospace too sadly.
Cheers!
int v =
((r << 24) & 0xff000000) |
((g << 16) & 0xff0000) |
((b << 8) & 0xff00) |
( a & 0xff) ;
Is there a reason you're doing the shift before the mask? I find this much less intuitive than the equivalent int v =
((r & 0xff) << 24) |
((g & 0xff) << 16) |
((b & 0xff) << 8) |
( a & 0xff) ;Cause it was like 1AM here when I wrote it?
Take your pick ;)
Yeah I'd totally write the bottom way I think, or I hope, cheers.
Have a look at Input [1]: a TrueType font built for coding, with somewhat-wide I and l and super-wide punctuation marks.
The issue is one of picking the right font, not of monospace vs. not - street signs, for example, have long used different fonts from printed pages, since they demand high accuracy and fast recognition of short chunks of text at long distances, as opposed to the pleasant long-reading experience demanded from most body fonts. Most systems just happen to have built-in variable-width fonts that are terrible for coding.
Another thing is easily formatting tabular-looking data (e.g. array initialisers) reasonably well. Proportional fonts mean this can't be done, and existing formatted data looks like a mess. Moving the cursor vertically with a proportional font also causes it to distractingly jump left and right instead of in a straight line.
Has anyone who prefers proportional fonts considered whether they'd like proportional vertical spacing (i.e. the height of each line varies depending on what characters are in it)? That would be the other extreme of variable spacing.
https://github.com/LuminosoInsight/python-ftfy/blob/master/f...
The result works pretty well in my gnome-terminal and Mac OS Terminal. You can probably see that GitHub in your web browser doesn't even come close to lining them up, though.
And the problem can't be completely solved, because the standards have gaps, and there are some scripts where monospacing just isn't a thing.
Chrome on Chrome OS (using FreeType 2) properly aligns the text in that <pre>, as does firefox on GNU/Linux (also using FreeType 2, in addition to graphite). On FreeType with Chrome or Graphite, the full-width latin characters are also rendered with the correct weight and face. Something that Chrome and IE on Windows seem to get wrong in your screenshots.
the rendered text for me on firefox (actually, not even proper firefox, but iceweasel, on debian stable, which is almost a year behind firefox) align perfectly fine. but the picture labeled "firefox" is all out of whack.
It's perhaps the difference between becoming a proficient Vim or Emacs user, and popping open Notepad.
To get back to Unicode, most of the hairiness could have been avoided if Chinese (and related regional variations) used a syllabic or alphabetic script - 100-200 characters vs ~80000 would make 2-byte wide chars sufficient to express nearly every non-dead script, with plenty of space for any number of poo or banana emoticons (if such a thing is really necessary, about which I have my doubts).
If you want to be cheeky you can argue that "hidden assumptions going from a 16 bit wide character set to a 20 bit wide character set" is still correct, because those hidden assumptions are part of UTF-16, and UTF-16 uses exactly 20 bits to express the astral planes, encoding them in a different way from the BMP.
Not sure how this'll come through on ycombinator, so it's also on a pastebin for 24 hours: https://paste.debian.net/311370/
~$ cd src
~/src$ mkdir œuvre
~/src$ cd œuvre
~/src/œuvre$ printf '#include <stdio.h>\nint main() {puts("hello œuvre");}\n' > œuvre.c
~/src/œuvre$ g++ œuvre.c -o œuvre
~/src/œuvre$ ./œuvre
hello œuvre
$ g++ --version | head -n 1
g++ (Gentoo Hardened 4.9.3 p1.1, pie-0.6.2) 4.9.3
$ uname -r -m
4.1.6-hardened i686
$ echo $LANG
en_US.UTF-8
$ bash --version | head -n 1
GNU bash, version 4.3.42(1)-release (i686-pc-linux-gnu)Naturally there's still the issue of properly identifying double-wide characters. But as long as you can correctly identify them (and for the ambiguous characters it generally treats them all the same, either as narrow or as wide depending on the software in question), you can simply render them with 2 cells instead of 1, and everything remains lined up.
But the article here was showing monospaced text in a web browser, and a web browser doesn't do any of this (nor a rich text editor). Font fallback usually attempts to maintain properties like monospacing, but the fallback font may still have different metrics, meaning it won't line up with the monospaced text from the source font (and depending on the fonts available, if you have no monospaced font with asian glyphs it would have to fall back to a non-monospaced font).
I suspect what's going on with the locale stuff here is simply that the locale affects the fallback list, and it ends up picking a different font that happens to have the right metrics to line up. But I can't say for certain that this is the explanation.
I somehow suspect those aren't the same people.