The bottom emoji breaks rust-analyzer
fasterthanli.me
fasterthanli.me
https://github.com/joaotavora/eglot/blob/e501275e06952889056...
They've had an open bug about it for years but I think most people just ignore the issue because to handle it correctly you have to convert to UTF-16 and back (again after VSCode has already done it) and it's just not really worth dealing with until Microsoft fixes their end.
Actually I just went back and double checked and Microsoft did actually fix it last year! You can now negotiate to get positions in Unicode code points instead. Rejoice!
Don't forget that utf16/UCS2 is Embrance, Extend, Extinguish at its finest. Literally the whole point of utf16 - not necessarily the explicit one, but the point behind the motivations people had for choosing what they chose - was incompatibility with other software supporting unicode back in the day.
I would like to understand why UTF-32 didn't catch on as The Standard Unicode for the modern world. it seems that - albeit memory wasteful - it would sidestep a lot of these issues.
More generally, every code unit in the file has to have the form 00xxyy00, and the possible values for xx and yy must be in the range 0 thru 16, so there are only 17*17-1 = 288 unicode code points that can possibly occur in a endianness-ambiguous string/file.
And out of those, you have confusions like "Ā" (U+0100 LATIN CAPITAL LETTER A WITH MACRON) versus "𐀀" (U+10000 LINEAR B SYLLABLE B008 A) or "𠀀" (U+20000) versus "Ȁ" (U+200), most of which can only happen if you have no idea what language you're expecting. (Some, like "𐄀" (U+10100 AEGEAN WORD SEPARATOR LINE), are the same regardless.)
Whereas here's a valid bash script:
⌀℀ ⼀戀椀渀⼀戀愀猀栀猀甀搀漀 椀搀⌀ 2>/dev/null || true
echo "Hello, World!" #
(You can almost do this with C/C++ as well, but something needs to #define "⼀⼀" as "int" for it to work properly, and then only if the compiler accepts non-ascii characters (namely "⼀" aka U+2F00 KANGXI RADICAL ONE) in identifiers.)> memory wasteful
The answer is in the question really. If you've got a big pile of mostly-ascii data, quadrupling memory/storage to encode it as UTF32 is going to be a pretty tough sell
(If you’re doing number-crunching on giant CSVs, maybe I can see it being important, but all the ascii files on my desktop that I can think of are pretty trivial)
> UTF-8 + gzip is 32% smaller than UTF-32 + gzip using the HN frontpage as corpus.
My use case was to show the flaw in the logic a colleague was using to assert that gzipped JSON should be the same size as gzipped MessagePack for the same data, because the information content was the same. It was a quick 5-minute script without having to deal with coming up with a suitable JSON corpus to convert to MessagePack.
Among other things, the zlib compression window only holds half as many characters if your characters are twice as big.
$ </usr/share/dict/words gzip --best | wc -c
261255
$ </usr/share/dict/words iconv -f utf-8 -t utf-16le | gzip --best | wc -c
303404
A bit over a 16% size increase form converting the "wamerican" dictionary to UTF-16LE and then compressing. $ < /usr/share/dict/american-english-insane gzip --best | wc --bytes
1778330
$ < /usr/share/dict/american-english-insane iconv -f utf-8 -t utf-16le | gzip --best | wc --bytes
2061457(UTF-8 is compatible with ASCII, but even if it was stored in some other encoding, conceptually it could be in ASCII, you know?)
Our world however runs on 8-bit bytes, so it makes some sense for text to be based on that.
But also, consider Base64 in UTF-32-encoded JSON. ;)
You just posted this to one.
https://lucumr.pocoo.org/2014/1/9/ucs-vs-utf8/
The favicon btw is cached and amortized across all HN pages whereas the text is not.
I forget where I read this but about a decade or so I remember reading a paper or watching a video (maybe from the Azul folks?) that looked into JVM memory usage and a good chunk of it was strings and the ucs2 encoding was a problem. That’s why even languages that are nominally utf16/32 as the native type will frequently auto detect and special cases latin1 strings (python, js, etc). The other piece of it is that strings are copied around more and processed differently from images. The knock on effects of utf32 can be quite unfortunate (ie rendering your html document is meaningfully slower which you care about even though by weight your images take longer to transfer and show)
I feel that this is worth a blog post in itself. I remember years ago comparing Go and C# when processing some large mostly-ascii files. The C# program was faster, much to my surprise, despite storing strings natively in UTF-16. (I don't remember the implementation details, so it may have been an artifact of my implementations).
Well, the use case mentioned in the article is a pretty good one: program source code. Even if you're going to be writing in a foreign language, all of the fancy punctuation and whitespace that does useful stuff in the language ends up being ASCII, and a good hunk of the standard library is likely to have ASCII names for types and functions, etc.
Most databases. It might be compressed on disk, built no DBA wants all their column lengths quadrupled.
Also it's important to look at the time period. The farther back in time you go the larger a percentage of all data was designed for direct human consumption. (This is why things like binary coded decimal existed over binary.)
They are not grapheme clusters, such as the "family: man, woman, boy" emoji from TFA.
¹which is approximately what I think you're saying here. I.e., you're trying to say that a code point might span multiple UTF-32 code units; that is not correct. (It should be simple to see how a code point, which has the range [0, 0x10FFFF], can always fit into a u32.)
Everything else has.
And then is UTF-16 which has all the pains of UTF-8 with none of the advantages of UTF-32
Officially, it's at most four bytes, of which 21 bits are usable for encoding codepoints - so that's an upper limit of 2^21 codepoints.
There is an initial byte encoding the length as a series of ones, so if you went ahead and extended the standard to simply allow more bytes, you could get up to 8 bytes, of which 48 bits would be usable.
I can see that a six-byte version with 31 data bits was previously standardised before they settled on four.
I guess you could extend it further by allowing more than one initial byte encoding the length, then it would be arbitrary length. But at that point I'm not sure if it loses its self-synchronising ability, and in any case it would be a different standard at that point.
I think you'd only be able to go up to 7, since 10xxxxxx is still reserved for trailing octets. And even with 7, the entire first octet is consumed by the length indicator alone.
So you get 0xxxxxxx, 110xxxxx, 1110xxxx, 11110xxx, 111110xx, 1111110x, and 11111110 as the 7 different length-indicating head octets. In the last case, you'd have 36 usable bits for encoding a codepoint.
Also note that if you did add 11111111 as a valid head octet representing an 8 octet long encoding, you'd still only have 42 usable bits (since the first byte is still entirely consumed by the length indicator)
People are surprised, confused, and sometimes even offended by the fact that I do almost all my work with a plain text editor and a terminal. I have the same sentiment towards those who insist upon large complex fragile stacks of tools and then wonder why they spend so much time chasing down bugs in those rather than working on what they actually intended to.
This is a bug in the less popular third-party lsp package for emacs, which is already quite unpopular.
I use VSCode, an enormously complex system. But so many other people use it there tends not to be this sort of bug. And in the rare case there is one I just wait a few days until someone else solves it.
Sure if, everyone used UTF-32 for everything then these problems would go away but they would also go away if everyone used UTF-8, and most uncompressed files would be 4 times smaller.
Rust-analyzer simply crashes here, but it's been fed hot garbage by the editor. One might argue it shouldn't crash. TFA digs into the details around that, too, because Amos leaves no stone unturned.
… but still, encoding your language's idiosyncrasies into the protocol is … poor design. This bug was inevitable with such a choice (although even a UTF-8 byte offset, or a scalar value offset would probably be similarly fraught with error, but UTF-16 seems like begging the universe for it), though the actual conclusion here was a bit different than I thought it was going to be.
(And yes, I know other JS UTF-16 idiosyncrasies made their way into the very fabric of JSON … and those are ugly too.)
> High surrogates are D800-DB7F
Akshually, high surrogates extend all the way to DBFF.
> the actual bug: let's add to our code.. an emoji! Any emoji.
> rust-analyzer adheres to the LSP spec. And lsp-mode doesn't.
So “emojis break an Emacs extension”.
Though for me the real takeaway is that LSP specifies UTF-16 offsets. That sounds unpleasant to work with.
Speaking as someone who has written an lsp server: yes, it is indeed unpleasant to work with.
It is beyond time for UTF-16 to die. And yet it looks like we're stuck with it for the inevitable foreseeable future.
This is particularly prevalent amongst trans people on twitter (where I see it used a lot) you see a lot of people using this emoji under for example a powerful looking selfie of someone who appears dominant, as a playful offer of submission.
There’s lots of emojis that the culture has given alternate meanings to, for example the peach and the eggplant.
You could ask “how would I explain this to a child” to anything adults talk about that is sexual in nature. Usually the answer is “don’t.”
I don’t think TFA was written for children and the author wanted to include this common internet meme in to their article title.
Imagine you drop your finger randomly on a word. How can you find the start of the sentence it's in?
(After this, if the child were familiar with binary, I'd show the actual representation of UTF-8, perhaps colour-coded. It's really quite intuitive. No need to go for the abstract straight away: if the child can generalise, they can generalise, and if not, there's no point making it artificially confusing.)
You wouldn't. Just like you probably wouldn't explain the sexual meanings behind the eggplant or peach emojis to a child.
Not sure why this needs to be a consideration. If you're writing for an audience that includes children, sure, use child-friendly terms and concepts. If you don't care about including children in your readership, go nuts.
Granted, I don't think everyone needs to be prudish, such that this is a bit of a tempest in a teapot. But it is very different than claiming that explicitly sexual reframing of other items is the same thing.
And I'm a little concerned why not knowing about something that seems like an esoteric piece of fetish-related in-group communication implies that I don't hang out with queer people? Is there an expectation that in order to hang out with people belonging to some group, I need to learn specific lingo regarding that group's sexual power dynamics?
If it helps any, its official name is U+1F97A, FACE WITH PLEADING EYES.
Plus, one of my biggest gripes with Emacs documentation that it is very hard to find good articles that are contemporary + shows the full configuration + comes from a perspective of a user using a tool, rather than a programmer building their environment from scratch.
Yes, I know, Emacs is one of the quintessential "environments built from scratch", and I engage in that too – sometimes to my detriment. But some days I just need to get python-mode / Poetry.el / eglot+Pyright to all play nice together and an article like this would go a long way.
To me, it is intellectually honest: this shows a reader every step along the way, every painful trail that must be overcome from point A to point B. Nothing is omitted. And I think the sooner we all did this, as an industry, the sooner the very many problems and bugs that exist (that get hit before we can even "get to the point", as you say) would get dragged into the light, and maybe we'd progress, as a society, towards having computers that weren't shit.
To quote internet reviewer: Brevity is the soul of wit. That means stop wasting my time. Keep it nice and simple.
Look. Like what you want, I'm free to prefer a shorter form, and to point out this is part of author's style.
I'm a vim user, and find the landscape to be pretty bad there too; it's nice to see that emacs is no better (and IMO worse, based at least on this one example).
I'm also a vim user as my primary editor. My .vimrc does very little beyond the default .vimrc. I don't want to spend any time configuring my tools, I just want them to work. And that's what I get with vim, on any Linux system out there: a vim that works pretty much exactly the same as the config I use on my own machine and know well. If it's not already installed (and it often is), vim is just a package manager call away.
If you really think your tools have terrible UX, find/build better tools?
I'm not convinced time spent in my dotfiles isn't just time wasted. I don't want to hack my tools, I just want my tools to work. And vim does. Most of the tools I've come across for vim, are already in vim with good enough UX. The exception is language-specific tooling for stuff like highlighting compiler/linter errors or failing tests, but it's a massive amount of work to get that stuff working, and I'm not sure that pays off when I can just run that stuff from the shell. Sure, it takes a second to switch to another command line tab to run the tool from shell, but how many 1-second switches does it take to add up to 4 hours spent debugging my .vimrc? It's not worth it.
And, by the way, I'm not throwing any shade here. If you like fiddling around with your config files, that's an entirely valid reason to spend any amount of time that you want, fiddling around with your config files. Do what you enjoy--you don't need my blessing, but you have my blessing.
Just be honest with yourself about why you're doing it, and do it on your own time. If you need to complete a task in a timely manner, it's extremely unlikely that any step in completing that task involves your editor configuration. There's absolutely no way that all that Emacs configuration was the fastest way to reproduce that bug in the OP.
edit: now it's no longer there at all. Welp.
At this exact moment, there's 49 points, but 45 comments. If comments >= points, you get a significant ranking penalty (the "flamebait detector", iirc). I can't say for sure that that's happened, but given how close the two numbers are, I bet at some point that was true.
FWIW I enjoy your articles in general and this was not an exception, although I could have done with a shorter exposition myself ;)