Unicode 16 now includes retro video game sprites [pdf]
unicode.org
unicode.org
Stop adding random emoji. Don't add fictional languages, no matter how cool. Don't do.. this.
By continuing to extend Unicode like this, they risk diluting their core purpose and creating unnecessary complexity. Unicode should remain focused on its original goal and not cater to niche or novel additions.
EDIT: I'm certain there's a proposal submitter out there that contorted the argument beyond the reasonable point that these existed in an old Usborne programming book text as inline images and need to be represented. I'm going to try to hunt it down.
EDIT 2: The original proposal? https://www.unicode.org/L2/L2019/19025-terminals-prop.pdf Some of the proposed characters might have come from an older submission as well https://unicode.org/L2/L2021/21234-terminals-smalltalk.pdf
Reply to EDIT: People who made the proposals are already well known in Unicode and they know what they are proposing. Michael Everson for example is responsible for encoding many other writing systems other than such symbols.
I guess I'll upgrade
Speak for yourself, I need this now!
In general, I think new emoji is going to be more persuasive to the average person to update than 'bugfixes to VCard deserialisation in the Contacts app' or the other usual release notes.
There's a lot of discussion about this; practically changing the execution character set isn't trivial, and it's not necessarily even a desirable feature. Even with full -fexec-charset support, it still makes sense to provide compile-time translation from Unicode to target strings. For example, Commodore PETSCII makes vastly more sense as an execution character set than a source character set.
I think it should be unnecessary to convert the character set; the execution shouldn't care about the character set except for ASCII and for whatever you program yourself in what the specific program you are writing is doing. It should not need to be a subset of Unicode, either; you should be able to use any character set that is a superset of ASCII (and where bytes in the ASCII range always mean ASCII characters and bytes not in the ASCII range always mean non-ASCII characters) (UTF-8 has this property and therefore may be used, but it is not the only character encoding with this property).
The C preprocessor is limited in its capabilities, although it would be helpful to add extra steps both before and after the preprocessor runs, which can transform character encodings, but also can be useful for other purposes too. (With GCC, I think this could be done by -no-integrated-cpp and -wrapper; I don't know about doing with Clang.)
(GCC will convert input to UTF-8 during preprocessing, but at least with the version of GCC that I have does not actually care if it is valid UTF-8 (at least for C; maybe not for C++ but I have not tried it), which is fortunate, since this means that you can implement your own character code handling.)
In the case of C++, as described there, you can use user-defined literals. They shouldn't require user-defined literals to be UTF-8 (nor Unicode), although if you can do whatever calculation you want on them at compile-time, then you can treat them as UTF-8 if you want to, but shouldn't be required to do so. (Personally, I do not use C++, though; so I do not actually know all of the details about how it is working, so I may have made a mistake.)
(There are several reasons you might deliberately not want UTF-8. One of them is security issues with the complicated text rendering involved with Unicode. Another might be the way that character widths are working. And there are many other possibilities, too. You might also prefer to put all non-ASCII text in a separate file; the #embed command can be used if you want to embed it into the program anyways, I suppose.)
> Even with full -fexec-charset support, it still makes sense to provide compile-time translation from Unicode to target strings.
Maybe, but I should think that this compile-time translation should be done separately as described above, and to be programmable to not be limited to only Unicode. It should not be required; I think it would be sensible that by default it should just pass through directly without conversion regardless of what the character set is.
> For example, Commodore PETSCII makes vastly more sense as an execution character set than a source character set.
I agree, but that is because Commodore PETSCII is not a superset of ASCII which is encoded as a superset of ASCII. The reason for this has nothing to do with Unicode.
I was thinking about how Minecraft has a system of components and layers that let you compose various flags on their banners. Obviously that's far, far simpler than country (and autonomous region, and county, and province, and and and) flags that can include text, symbols, and practically entire images. But I did wonder if there was some way that could be represented. Unfortunately, I'm not nearly well-versed enough in code points and their ilk to propose anything useful.
But, I am torn. Archival projects are important, too, and language evolves. These decisions will live for potentially hundreds or thousands (Linear A) of years, and interoperability in computing is important.
> Identities are fluid and unstoppable which makes mapping them to a formal unchanging universal character set incompatible.
If you really want to have identity flags encoded in spite of that, you don't really need Unicode's blessing. The pride flag is already not a single character anyway, it's U+1F3F3 U+FE0F U+200D U+1F308 (white flag -- emoji force -- ZWJ -- rainbow) and you can always create new ZWJ sequences with your own font. Or you can make a font that automatically synthesizes flags from some ZWJ sequence pattern, which is no longer semantically valid but should be much more flexible. Once they got sufficiently popular, there is no other reason that your new ZWJ sequence(s) shouldn't go into Unicode per se.
Unicode's decision to not process non-country flag emoji proposals is because they are closely tied with (minority) groups and Unicode wasn't expected to do any resulting conflict resolution. If you can somehow resolve that problem in advance, then you should probably do that first and propose what you've done.
Some of them clearly look like copyrighted characters from specific games, space invaders and PAC MAN in particular.
So I'm thinking it's more for a snake game sprite than pac man. With a solid block for the body, the "tail" to connect the head so you don't have a floating head.
You could conceivably combine them to create any sprite
[1] https://www.unicode.org/L2/L2021/21234-terminals-smalltalk.p...
[2] https://www.unicode.org/L2/L2021/21235r-terminals-supplement...
The alternative to "a new version of Unicode every year" is not stasis, but rather new and incompatible encoding schemes frothing like JavaScript frameworks.
IIRC, it was initially limited to 16 bits outright.
The Unicode Standard, Version 16.0
Would be an interesting project
https://en.wikipedia.org/wiki/Cyrillic_O_variants#/media/Fil...