Unicode 18.0.0 Beta
unicode.org
unicode.org
Emoji proposals and status: https://unicode.org/emoji/emoji-proposals-status.html
To elaborate: it should be plain obvious that not every Emoji proposal can be accepted even though all of them are correctly filed, as there would be too many Emojis there then. So there has to be some threshold, and that threshold is mostly stipulated by vendors' willingness to process new Emoji characters for designing fonts and updating softwares in time.
Since it's a vote, there is no single official 'reason' for rejection. If I had to guess: it would be confusing to anyone who didn't grow up with American TV shows.
I don’t protest the coinage here (goodness knows my native language did worse things to English words), but I can’t stop saying it in Gollum’s voice.
it's a popular image/byword/archetype for conspiracy theorists, idk if it's a common enough symbol to justify emoji inclusion. the submitted proposals probably have analyses of that though :p
So you generally can’t add something because it would be cool or fun or useful, but only because it is currently in use and cannot be encoded by Unicode.
Consider all of the languages it supports. Consider: ﷽ (which isn't an emoji, but the point stands) which is an entire sentence. It was already in use in certain places and unicode decided they wanted to support it, so now they do. Previously, one would have to type out the entire sentence in the original characters, but now it is a single unicode, just like u+263a () used to be alt+1 (). The emoji was already in use long before unicode existed, and in seeing it in common use, they decided to support it.
For the most part, now that Matrix is merging those Matrix 2.0 specs finally, and the 2.0 features are already out in the wild with excellent results, it has a really good base, and as expected we've started to see clients build more into the average-consumer space to pose as alternatives to both niche and mainstream audiences such as Discord, Whatsapp, etc - Which it just wasn't/isn't able to do on Matrix 1.x (legacy).
Commet has an open PR for this but not yet implemented.
- is there an usable font the cover all unicode ?
- if not is there really a point to include everything possible in unicode ?
- how many space is remaining for new alphabet and smileys ?
- how do they handle changes in scripts, for example if new proto-cuneiform or seal script symbols are discovered ?
There is also GNU unifont [1] "The original intent of Unifont was to offer a simple font format with wide Unicode coverage to render something meaningful for each Unicode code point"
Needing to load three fonts to show a single document that mixes vastly different character sets is still infinitely better than not being able to have those different characters in the same .txt or .md file at all
> how many space is remaining for new alphabet and smileys ?
Unicode can encode about 1100k code points, and about 800k of those are currently unassigned and available for future scripts or characters
They get added in the next Unicode revision.
In Unicode you have "blocks" [0] that are often bigger than the number of characters in a script, language or function. There are usually also space for new blocks between unrelated blocks.
For example, in the case of cuneiform, it was introduced in Unicode 5.0, and there have been revisions in 7.0 and 8.0 [1]
--
0: https://en.wikipedia.org/wiki/Unicode_block
1: https://en.wikipedia.org/wiki/Cuneiform_(Unicode_block)#History[2] https://www.unicode.org/L2/L2025/25017-miscellaneous-musical...
* Cracking face
* Left/Right thumb sign
* Monarch butterfly
* Pickle
* Lighthouse
* Meteor
* Eraser
* Net with handle
I mean, making or help making sovereign AI models is nowhere near responsibilities of Unicode, but Han Unification and sort of a default-enforced IVD support is literally adding small but non-zero amount of fuel to cultural division and xenophobia perpetuate in East Asia. I doubt blaming users would work here.
> cultural division and xenophobia perpetuate in East Asia
By the way, I recently have seen multiple claims from Japanese Twitter users that Korea would have been better keeping Chinese characters (Hanja) in use. If this is a cultural division and xenophobia we are talking about, I will gladly take it---why on earth do they have any saying in Korea's choice of scripts? The "sinosphere" is an illusion, the fact that CJKV countries have or had shared the same set of characters is just a fun fact and not a cultural mandate or anything else like that.
Maybe, but no one is running an ivdfy-filter through every single Japanese documents and the issue keeps going. Maybe one way to make it happen is to make the Simplified forms singularly canonical to the CJK Unified Ideographs so to classify everything in that form as Chinese, and define Japanese script as being always flagged with IVDs, though I don't know what the storage and processing implication of that might be. But my point is that maintaining the position that users can optionally choose to not display text in a wrong language and Unification issues are merely user errors don't make any sense to me.
> Korea would have been better keeping Chinese characters (Hanja) in use.
I can't speak for all, but I, for one, do regularly encounter machine translation failures in Korean contents due to homophones even with LLM-based ones in the ways that don't happen with Japanese. It manifests as either homonym errors[1] or the MTL resorting to phonetic transcripts that I have no idea about[2]. Both happens in formal writings like newspaper Web articles in addition to casual social media posts. Since it appears that there's no way this issue could happen with "our" system, it sometimes feel like reverting to that could fix it.
1: (like "plain/plane", had the source been English and this was somehow happening)
2: (like "That arm might be fukuzatukossetsushiteru" had the source been Japanese)
Unicode takes backward compatibility like this very seriously.
- Left and Right parenthesis with middle ring [1]
- A wiggly exclamation mark expressing mirth or laughter [1] (edit: and something I completely missed: the inverted version can express sarcasm)
- Cuneiform numerals, including lots of arranged dots that might be useful in other contexts [2]
- New variations of "measured angle" and "sector" [3]
- A transparent cube and a white cube [4]
Also a couple new combining marks
And for anyone who wants to see what the reference images for the new emojis look like:
Lighthouse: https://www.unicode.org/charts/PDF/Unicode-18.0/U180-1F680.p...
Other new Emojis: https://www.unicode.org/charts/PDF/Unicode-18.0/U180-1FA70.p...
1: https://www.unicode.org/charts/PDF/Unicode-18.0/U180-2E00.pd...
2: https://www.unicode.org/charts/PDF/Unicode-18.0/U180-12550.p...
3: https://www.unicode.org/charts/PDF/Unicode-18.0/U180-1CEC0.p...
4: https://www.unicode.org/charts/PDF/Unicode-18.0/U180-1F780.p...
https://www.emilydamstra.com/please-enough-dead-butterflies/
13615 towctrans-5.h
14889 towctrans-6.h
15477 towctrans-7.h
15815 towctrans-8.h
16454 towctrans-9.h
16460 towctrans-10.h
16756 towctrans-11.h
16955 towctrans-12.h
17068 towctrans-13.h
17456 towctrans-15.h
17456 towctrans-14.h
17701 towctrans-17.h
17721 towctrans-16.h
18620 towctrans-18.h
Made with cpan Unicode::Towctrans;
for v in `seq 5 18`; do gen_wctrans --out towctrans.h -v $v --ud UnicodeData.$v.txt; doneThese are already optimized folding-tables.
At least nothing is wiggling. Of those Unicode points which are graphical, at least all of them can still be printed on paper and won't require a screen. I wonder how long that invariant lasts.
Also, in passwords on websites to keep developers on their toes.
And yeah, :slack-style-emoji-notation: is superior. It was just a historical necessity for Google/Apple.