You mean like language-dependent unified Han glyphs? [1]
[1] https://en.wikipedia.org/wiki/Han_unification#Examples_of_la...
Animated emojis probably aren't too far off the horizon, and I could definitely see a world where multiple glyphs could be combined to affect the animation much like we do for skin tone or the family emojis.
Sometimes I think we've gone too far.
Agreed from a technical standpoint. But also dont forget we have the poop emoji and the love-hotel.
Now what I find is less practical and more ridiculous is the sheer number of ways you can combine various codepoints to get different emoji representing families. On one hand, it’s impressive dedication to trying to be neutral and encompassing, on the other hand it is nightmarish just how many possibilities exist.
Another example is the omission of the middle finger emoji in earlier Unicode versions but inclusion of various different hand emojis that insult people in East Asia or the Middle East. E.g. OK Hand Sign and Call me hand, which are both completely mislabeled when used in other cultures.
My problem with this is that it's trying to solve a font-level problem at the alphabet-level. For example, it's fair to ask why a male-looking construction worker appears on some particular emoji keyboard, whilst a female-looking construction worker doesn't; yet that bias exists in the choice of font, not in the Unicode standard (where U+1F477 simply defines 'construction worker').
Vendors like Apple created this problem when they moved away from simple silhouettes and (AFAIK) neutral 'smiley faces', to more detailed images which required a bunch of arbitrary choices to be made (gender, skin tone, etc.). Rather than making a bunch of alternative fonts, akin to light/bold/serif/monospace/etc. those choices were instead shoehorned into Unicode's modifier-character system :(
identifier -- printable ASCII characters only, an array of 8-bit chars
ucs16 -- An array of 16-bit chars for compatibility with Windows, Java, and .NET
utf8 -- Normalised, 100% valid UTF-8 with potentially some "reasonableness" constraints
text -- Abstraction over arbitrary code pages, including both Unicode and legacy encodings.
Languages like Rust kinda-sorta implement this. For example, the PathBuf type internally uses a "WTF8" encoding that is vaguely Utf-8 compatible, but allows the invalid code sequences that can turn up in Win32 system API calls.IMHO that's a good try but not ideal. The back-and-forth conversion is complex, requires temporary buffers, and is slow as molasses for many types of API calls.
The ideal would be to have abstractions (traits, interfaces, whatever) that cover all the use-cases. E.g.: it should be possible to test if a 'utf8' string contains an 'identifier' string. It should be possible to compare strings without having to convert their formats. Etc...
# Simple indexed access (often an array, possibly an array of arrays)
raw / octets -- Not-classified sequence of raw bytes
# Fancy strings, which MAY be validated (but don't have to be), and MAY stay validated (if the operations are known to be simple enough), and MAY also have multiple types of index for speedy access to specific points by raw byte, unit run of encoded components, complete display units (a single displayed element), or even a cached last known render left bound on an output.
UTF-8 / WTF8 / UTF8 / ASCII -- A fancy string with octet components of possibly multi-byte (variable) length
UCS-2 / USC-16 / etc -- A legacy string encoding format that no one should use as it too is variable length but suffers from endien confusion in raw data storage.
In an object based language the latter two would probably use the first as a raw storage mechanism for the strings, while they'd also have some associated attributes for the desired encoding, if it's validated, how it is known to be normalized, and storage for different index aids.
Crucially the system libraries should have the same interface. If there isn't library support for converting / normalizing encoded text the results should always be raw octets. If library support is included then WTF8 should probably be the result target of any operations, possibly upgraded to UTF-8 if the results are theoretically still valid.
Just no. Maintaining compatibility with broken implementations that hold incorrect assumptions is not something that should be done. Especially not encoding it into some kind of standard.
Those three platforms describe 90-95% of all "enterprise" business software ever written or the platform they're used on.
Disregarding that weight of history for... what? A clever trick that Linux used to finally add i18n support decades after other platforms?