Terminal support for emoji
darrenburns.net
darrenburns.net
It's not the case here, but for some emoji there's another issue: Unicode 9 changed the width for some codepoints (mostly emoji) from 1 to 2, and iTerm until very recently (don't know if it's released yet) defaulted to the Unicode 8 widths, with an opt-in escape sequence to change to Unicode 9.
>This approach comes from the wcwidth utility, and the comment at the top of the C source file provides further insight into the difficulties faced here.
That's link goes to Markus Kuhn's implementation from 2007. It supports Unicode 5, and is by now woefully out of date. You don't want to use it anymore.
Most terminals have their own definition, and the annoying part is that the client application and the terminal need to have theirs in sync or they get weird glitches when moving the cursor.
Shameless plug: Fish's solution is widecharwidth[0], which is a python script that parses the Unicode data files and generates a wcwidth for C++, Javascript and Rust. It's still a wcwidth, meaning that it has issues with joining code points, but it's at least a start. It's up-to-date with Unicode 14 and, unless they change the data format (again) should be easy to update to future Unicode releases.
It's public domain and used by at least fish and WezTerm.
Will do. If its more complicated than a box-character, it doesn't belong in my terminal :-)
Or that I can't type my address, which contains ø?
Poop emoji might be dumb, but it doesn’t seem to hurt anything and if it makes developers support Unicode better, awesome.
Clearly the article is demonstrating that this hasn't come true.
Take something as simple as names: https://www.kalzumeus.com/2010/06/17/falsehoods-programmers-...
That said, I have always assumed that these problems we get with emoji also have equivalents that show up in other languages, it's just that we speak English and so only see them with Emoji. Maybe it's good that Emoji were included because it actually stresses our Unicode implementations!
And this is all before you introduce text that is written right-to-left like arabic scripts.
Just my 2 cents, but if ASCII includes, among other things, "characters" for linefeed, 4 unspecified "device control characters", EOT (End of Transmission), and a character to ring a bell, then I guess its okay to put some smileys and flags into Unicode.
The problem is the massive quantity of LTR stuff that is in RTL languages still (English words, programming language keywords, math, music, certain conventions, ...)
Anyway yes multi-character glyphs occur in various forms (like accents, sometimes: é, ñ...) but I always thought that the burden of supporting them was on the fonts, not on the Terminal or the Unicode standard itself.
I am well aware, and supportive, of the need to encode the entirety of human script and written symbolics in a unified format.
And the reason why my terminal doesn't need anything beyond basic modifying diacritic marks and single-codepoint symbols (like blocks and line characters for text-user-interfaces) is because I do not use it to write prose is a variety of languages, I use it to communicate with a machine, administrating systems, designing backend software, and analysing machine-written logfiles.
I am NOT a native English speaker. I cannot even write my name with ASCII letters.
To me, the situation looks exactly the reverse of your comment, I see the urge of putting Unicode everywhere as an essentially white savior syndrome that nobody asked for.
I want to be able to write name in Word or other text processing software, but nothing more. The command line, network protocols, etc. There are not the realm of fancy dynamic length characters. I want simple glyphs with fixed size bytes mapping. I don't want to embed a complicated Unicode parsing library in every software.
When I'm in the realm of programming, I'm not here to do politics. I'm here for the most reasonable, straightforward, working solution for a set of problems. If that means accepting a fixed set of characters, I'm 100% fine with that.
The cat is out of the bag, and if your software doesn’t support Unicode. you’re limiting what your users can do with it. Maybe that’s fine with you and a certain program, but it’s not for a lot of people.
You also don’t need to embed full Unicode parsing/support libs if you just treat the user’s input as a byte stream. When programmers try to get fancy and ToUppercase() a byte blob with no regard for what that blob represents - that’s when we have problems.
So how does one write, say, the backend for a reservation system managing thousands of hotel room bookings every hour, all over the world? I'm pretty sure there will be ALOT of customers with non-ascii names, not to mention adresses and hotel names. How does one write a system like twitter, where messages in every known system of writing need to be be sent, processed, stored, delivered, commented, displayed?
And btw. network protocols and backends work with arbitrary data all the time: Audio & video streaming, networked gaming, sensor readouts, images. Even plaintext webpages are often transmitted in compressed form. So if these systems are the "realm" of arbitrary bytestreams to encode everything from classical music, over sha-hashes, to gifs of dancing dogs, they may as well encode a few hundred different systems of writing and some smiley faces.
Because lots of families don't look like that, and people want to use emoji that reflect them and theirs
Flip phone emojis started in mid-1990s as carrier-specific custom characters on unused areas on Shift-JIS, and initially the implementations were largely the same. Wireless carriers then realized it could be used as differentiators, like by adding more specific mojis, colored mojis and animated GIF emojis, and then each strains of emojis that carriers offer grew into different dinosaurs over a decade and half until forcibly unified into a common ground in 2010s basically by Apple and Google.
Apple got into the emoji game when they entered Japanese market with iPhone 3G on SoftBank, and Google then helped it standardized in Unicode. At that exact point everyone was on the same page. And at next instant, they realized that they are not differentiating on emoji.
Turning Unicode into clipart collection is not.
https://github.com/microsoft/terminal/issues?q=is%3Aissue+is...
Unfortunately, as far as I know there is no "complete" monospace font set that handles emoji. Ideally, you want a font with two character widths, with double-width for emoji, hanji (CJK characters), and similar. Instead the browser will substitute these characters from some variable-width font, and then the spacing will be off.
DomTerm handles this by putting double-width characters as well as Extended Grapheme Clusters in a separate span that is forced to have the correct width. I created a JavaScript library https://github.com/PerBothner/unicode-properties based on other people's code but optimized for DomTerm's needs: It provides both East Asian Width (for recognizing double-width characters) and character classes (for grapheme clusters) in a single efficient trie structure.
The DomTerm equivalent of tmux's "select mode" is grapheme-cluster-aware, so left/right-arrow will correctly move over an entire grapheme cluster.
Sounds like a bug which could be fixed.
But I have a thing for using built-in tools and Terminal.app gets a lot of the basics right. And has for so many years.
echo '\U0001F3F5\uFE0Fhello'One of the simplest workarounds is to ensure that the default emoji font is black&white; the linked issue above suggests other workarounds.
Also in the official st FAQ: https://git.suckless.org/st/file/FAQ.html
(search for "when trying to render emoji")
You can get the fancy one with print("\U0001F6E5\uFE0Fhello"). However, the bug that the post alludes to is that not all terminals override the east_asian_width to 'W' when they see "\uFE0F". For example on gnome-terminal the "h" overlaps with the right half of the boat.
worked in putty too
Here's a quick guide:
https://towardsdatascience.com/converting-emojis-to-text-a28...