So how far off is Unicode from being 'done'? At what point will they be able to stop adding characters and scripts?
So how far off is Unicode from being 'done'? At what point will they be able to stop adding characters and scripts?
The newly added letter is the "capital letter Eszett", which did not exist until recently. One could argue that this new letter is not really needed, as Eszett does not appear in capitalized form except when a word is in all-caps, and was then simply written as "SS".
SZ hasn't officially been an option at least since 1996.
ß is both visually and from its name (‘s’ ‘z’) a ligature of s and z (“Eszett” in English is “ess-zed). An ss ligature would look like a double integral sign. “Sz” seems like a better way to represent that sound, so I don’t know why it morphed into “ss”
Greek has two characters for s, one for use in the middle and one for the ends of words. English lost this in the late 18th or early 19th century (look in the Declaration of Independence for examples). German kept it longer, at least in Fraktur, which included other standard ligatures, even in handwritten text, such as ch, tz et al. The Umlaut mark can also be considered a ligature for E which is how it was originally drawn.
That letter, the capital eszett, has existed in German typefaces since at least 1905:
> Historical typefaces offering a capitalized eszett mostly date to the time between 1905 and 1930. The first known typefaces to include capital eszett were produced by the Schelter & Giesecke foundry in Leipzig, in 1905/06. Schelter & Giesecke at the time widely advocated the use of this type, but its use remained very limited.
Eszett is usually just a lower-case form; it is most often upper-cased to SS, or two capital letter esses, being an example of how changing case does not always preserve the number of letters in a text. Capital eszett is very rare, but it was in uncommon usage in German text, and so it was added to Unicode.
Do you think human written language is ‘done’ and will never evolve?
But you wait until Maya script gets into Unicode. You'll have at least three different Maya codepoints for death. (-:
Where do you draw the line? I draw it at "anything in or using a language that people might write in the absence of computers, which they would then reasonably want to store and transmit using a computer". I don't include "any possible visual communication that can occur using a computer". That's far too broad to define "text", or be part of any existing "language", which are the stated goals of Unicode.
Furthermore, I'm really starting to question the way CJK is encoded. We don't make every English word a separate codepoint. 97% of these CJK ideographs are just different combinations of the same few radicals. Korean seems especially weird, as they have both individual radicals and every precomposed triple (in a block that's been rearranged once or twice, on the basis that nobody was really using it yet). I'm not saying we should nix all precomposed Hanzi/Kanji, exactly, because that's a very convenient way for programs to handle text, but it seems like this system is becoming increasingly awkward for non-western languages.
I feel there's a fundamental flaw when our "universal" text encoding system can't handle the regular creation of new words in a well-understood way, for languages spoken by 1/3rd of the world's population. It's like we're issuing hardware patches for a software problem.
Even the "C" has traditional and simplified variant.
Fortunately I think Unicode is pretty much done for Alphabetical languages. Someday if CJK design Unicode isn't good enough breaking it off to something better isn't entire impossible.
Many of the additions today, barring emoji, are covering historical usage. This includes things like Medieval scribal annotations, a different set of numbers for the Ottomans, and the Mayan script. It will still be over a decade for the historical work to be complete, since there is often a lot of actual research that needs to be done to understand how an ancient writing system works, which has to come before you can even put together a coherent proposal for a new script.
* https://www.emojis.com/food/fruit/
There will always be another Emoji that someone, somewhere wants to add.
And while every Unicode announcement gets derided because they added more emoji, it's still just a small subset of the standard. Back when they added them I wasn't much of a fan, but by now I think it was the right decision. I still don't use them, but I've heard they're quite popular in younger age groups.
[0]: https://linguistics.berkeley.edu/sei/index.html [1]: https://linguistics.berkeley.edu/sei/scripts-not-encoded.htm...
I wonder why Gardiner's descriptions never made it into Unicode? https://en.wikipedia.org/wiki/List_of_Egyptian_hieroglyphs#E
O, wait, there are censored genitalia in D block ...
My wifi SSID is <horse U+1F40E><unicorn U+1F984>, which is a barely satisfactory approximation.
(FWIW the FTP site is still up: ftp://ftp.unicode.org/Public/13.0.0/ )
This part is dear to me, as I helped craft it. It includes 2x3 videotext mosaic characters that will make it much easier to draw large text have have better quality charts in text terminal interfaces.
And, of course, the ability to properly encode documents that were generated in computers in the 70's and 80's that contained those platform-specific characters.
For 14 we are planning on adding symbols from the Sharp MZ series and the large text characters (3x3 cells) of HP terminals.
Found a unicode consortium tweet (!) with a picture of the whole block:
https://twitter.com/unicode/status/1085613123183071232?lang=...
* http://pelulamu.net/unscii/ (https://news.ycombinator.com/item?id=18478350)