Unicode 13.0
unicode.org
unicode.org
> Symbols added in this release (which aren't implemented as emojis) include a Creative Commons symbol, as well as other related glyphs for non-commercial licences, or to indicate where attribution is required.
So you can now indicate license requirements, such as attribution, share-alike, or non-commercial use only. Those can apply to a vast number of works, so does nice to have easily-accessed standard symbols for it.
It also adds symbols for some older machines, which will make discussing them easier.
And yes, it has new emojis. But it is worth noting that the total number of emojis is a very small portion of the Unicode characters.
(This is why there is no Apple logo or character for TAFKATAFKAP, for one thing.)
https://creativecommons.org/wp-content/uploads/2016/10/CC-Un...
The CC logo and icons as a matter of trademark law are governed by CC’s trademark policy, which specifies conditions of their use “in a manner reasonable to the medium and context.” According to Unicode criteria for encoding symbols, a trademark weakens, but does not disqualify, the case for encoding. As such, we would like to address this point directly, in addition to emphasizing the criteria that strengthens the case for encoding. So as to be sure there exists no misunderstanding, CC is seeking inclusion of the CC logo and icons in the Unicode standard notwithstanding that CC asserts trademark rights in its logo and icons.
The aims of Unicode and the CC trademark policy are not antagonistic to one another. It is perfectly consonant with the purposes for Unicode to allow trademarked logos and icons into its standard without jeopardizing trademark rights of the requesting organization. Especially where, as here, there is a clear and public trademark policy by the submitting organization in place that clarifies when and how the logo and icons may be used, and such marks are ubiquitous.
Though the CC logo and icons are trademarked, they are a widely-used functional marking tool that indicates permissive use as an alternative to the well-known © symbol, followed by the icons that represent those terms.
As such, they were accepted and celebrated by MoMA into its permanent collection,6 alongside universal icons such as the @ symbol and the International Symbol for Recycling—both of which are encoded in the Unicode standard. Encoding the CC logo and icons in UCS would more easily enable creators to mark their works as consistent with CC’s trademark reasonable to the medium and context, in this case within text-based editors.
I'm personally hyped by the new characters for Creative Common, that could come in handy in data processing (at least in my field).
It's cool the Bopomofo was extended to support Cantonese (didn't know this existed tbh), because it means multilingual Madanrin-Hokkien-Cantonese can use a uniform spelling for word and character readings.
And finally Khitan small script, while ghetto, was an interesting offspring of the Chinese script that targeted a Mongolian language.
PS: I won't comment on the ridiculous emoji shitshow.
So how far off is Unicode from being 'done'? At what point will they be able to stop adding characters and scripts?
I wonder why Gardiner's descriptions never made it into Unicode? https://en.wikipedia.org/wiki/List_of_Egyptian_hieroglyphs#E
O, wait, there are censored genitalia in D block ...
My wifi SSID is <horse U+1F40E><unicorn U+1F984>, which is a barely satisfactory approximation.
(FWIW the FTP site is still up: ftp://ftp.unicode.org/Public/13.0.0/ )
This part is dear to me, as I helped craft it. It includes 2x3 videotext mosaic characters that will make it much easier to draw large text have have better quality charts in text terminal interfaces.
And, of course, the ability to properly encode documents that were generated in computers in the 70's and 80's that contained those platform-specific characters.
For 14 we are planning on adding symbols from the Sharp MZ series and the large text characters (3x3 cells) of HP terminals.
Found a unicode consortium tweet (!) with a picture of the whole block:
https://twitter.com/unicode/status/1085613123183071232?lang=...
* http://pelulamu.net/unscii/ (https://news.ycombinator.com/item?id=18478350)
The newly added letter is the "capital letter Eszett", which did not exist until recently. One could argue that this new letter is not really needed, as Eszett does not appear in capitalized form except when a word is in all-caps, and was then simply written as "SS".
SZ hasn't officially been an option at least since 1996.
ß is both visually and from its name (‘s’ ‘z’) a ligature of s and z (“Eszett” in English is “ess-zed). An ss ligature would look like a double integral sign. “Sz” seems like a better way to represent that sound, so I don’t know why it morphed into “ss”
Greek has two characters for s, one for use in the middle and one for the ends of words. English lost this in the late 18th or early 19th century (look in the Declaration of Independence for examples). German kept it longer, at least in Fraktur, which included other standard ligatures, even in handwritten text, such as ch, tz et al. The Umlaut mark can also be considered a ligature for E which is how it was originally drawn.
That letter, the capital eszett, has existed in German typefaces since at least 1905:
> Historical typefaces offering a capitalized eszett mostly date to the time between 1905 and 1930. The first known typefaces to include capital eszett were produced by the Schelter & Giesecke foundry in Leipzig, in 1905/06. Schelter & Giesecke at the time widely advocated the use of this type, but its use remained very limited.
Eszett is usually just a lower-case form; it is most often upper-cased to SS, or two capital letter esses, being an example of how changing case does not always preserve the number of letters in a text. Capital eszett is very rare, but it was in uncommon usage in German text, and so it was added to Unicode.
Do you think human written language is ‘done’ and will never evolve?
But you wait until Maya script gets into Unicode. You'll have at least three different Maya codepoints for death. (-:
Where do you draw the line? I draw it at "anything in or using a language that people might write in the absence of computers, which they would then reasonably want to store and transmit using a computer". I don't include "any possible visual communication that can occur using a computer". That's far too broad to define "text", or be part of any existing "language", which are the stated goals of Unicode.
Furthermore, I'm really starting to question the way CJK is encoded. We don't make every English word a separate codepoint. 97% of these CJK ideographs are just different combinations of the same few radicals. Korean seems especially weird, as they have both individual radicals and every precomposed triple (in a block that's been rearranged once or twice, on the basis that nobody was really using it yet). I'm not saying we should nix all precomposed Hanzi/Kanji, exactly, because that's a very convenient way for programs to handle text, but it seems like this system is becoming increasingly awkward for non-western languages.
I feel there's a fundamental flaw when our "universal" text encoding system can't handle the regular creation of new words in a well-understood way, for languages spoken by 1/3rd of the world's population. It's like we're issuing hardware patches for a software problem.
Even the "C" has traditional and simplified variant.
Fortunately I think Unicode is pretty much done for Alphabetical languages. Someday if CJK design Unicode isn't good enough breaking it off to something better isn't entire impossible.
* https://www.emojis.com/food/fruit/
There will always be another Emoji that someone, somewhere wants to add.
And while every Unicode announcement gets derided because they added more emoji, it's still just a small subset of the standard. Back when they added them I wasn't much of a fan, but by now I think it was the right decision. I still don't use them, but I've heard they're quite popular in younger age groups.
Many of the additions today, barring emoji, are covering historical usage. This includes things like Medieval scribal annotations, a different set of numbers for the Ottomans, and the Mayan script. It will still be over a decade for the historical work to be complete, since there is often a lot of actual research that needs to be done to understand how an ancient writing system works, which has to come before you can even put together a coherent proposal for a new script.
[0]: https://linguistics.berkeley.edu/sei/index.html [1]: https://linguistics.berkeley.edu/sei/scripts-not-encoded.htm...
"Support for these legacy computing symbols includes 212 characters added in Version 13.0 to provide compatibility with a wide range of early home computers, or “microcomputers,” manufactured from the mid-1970s to the mid-1980s. These symbols also cover the teletext broadcasting standard originally developed in the early 1970s, and the Minitel standard developed in the 1980s. This collection of early microcomputer symbols includes support for the character sets of Amstrad CPC, Apple 8-bit, Atari 8 and 16- bit, Commodore 8 and 16-bit, MSX, Yamaha, RISC OS, and Tandy"
I need to check my combining class stuff, it seems.
For a long time on its homepage placed message banner:
> New versions of all fonts will be posted the release of Unicode 13, March 2020.
I predict that within 10 years, some group will start a new simplified text encoding standard which is just for text.
On the other hand, that'd be dependent on connectivity.
[0] https://blog.leahhanson.us/post/recursecenter2016/haiku_icon...
Course I find it silly that they swapped a real gun emoji for a water gun. How will my friends know I intend to go hunting and not play water guns with my nephews or whatever.
That is not the issue. The issue is that users expect the emoji they send to look the same to the receiver.
Geeks know that emojis are "characters" and therefor might look slightly different in different fonts. But this is not how the average user sees it. They see emojis as small images and expect them to transfer unaltered just like any other image they put into a message.
When they were really simple icons, it made sense to treat them as characters. But today they are detailed illustrations.