Banish missing glyphs with Unifont
shkspr.mobi
shkspr.mobi
And in fact, the font actually looks pretty good under a CLI console. I once tried a Linux kernel patch that embedded the 10 MiB+ Unifont to the native Linux VT console, and I got a perfect multilingual console (without using a 3rd-party framebuffer-based or KMS-based console), including all the CJK characters. It's also a good choice for a dot-matrix LCD/LED display.
No, you didn't. If it was using Unifont, there are many scripts and languages that it was unable to display correctly because they require "shaping" that Unifont cannot support; it lacks the necessary glyphs, let alone the OpenType (or equivalent) tables to control the rendering.
Displaying one nominal glyph per Unicode character does not result in readable text for many languages. Just because it seems adequate for Western alphabets or for CJK doesn't mean it is sufficient for true multilingual support.
Many millions (billions?) of people won't be able to use your terminal in their language.
Completely unreadable, even to people with a basic understanding of the inner workings of Unicode and some practice with the font, or just not matching the norms of correctness?
Decipherable but clearly not correct would mark the sweet spot between being just as bad as placeholders and being so good that it undermines efforts for having a proper font for a given language.
Can you give an example?
So "spelled-out", क्षत्रिय is made up of क ् ष त ् र ि य . I don't know if Unifont tries to render the combining characters as zero-width overstrikes, or what, but much more than that is needed for a readable result.
(This comment is a good equivalent of the problem in English. https://news.ycombinator.com/item?id=19575942 )
It does look different (likely incorrect), but is definitely better than blank boxes and question marks.
I disagree, I think it's better for it to be apparent when text is wrong or unreadable, even to people not literate in the language.
I wondered if the problem had to do with Unifont being monospaced but I pasted "क्षत्रिय" in on the Inconsolata page and it appears correct (i.e. the lines correspond to how they look in the font on this page).
(IMO, the fact that fontlibrary.org doesn't indicate this is a serious defect in its presentation.)
Arabic has a strong difference between letters of the same word, and different words. There is no exact equivalent in English...
b&u&t&i&t&m&i&g&h&t&b&e&s&o&m&e&t&h&i&n&g&l&i&k&e&t&h&i&s
b&t&u&t&i&m&g&h&t&i&b&s&e&m&o&t&e&h&n&g&i&l&k&i&t&e&h&s&i
* http://jdebp.uk./Softwares/nosh/user-vt-screenshots.html#Uni...
* http://jdebp.uk./Softwares/nosh/guide/terminal-resources.htm...
Your multilingual console is really far from "perfect". Aside from the aesthetics of GNU Unifont, you do not have good multilingual input, which you have completely left out. You don't have an ISO/IEC 9995 common secondary group. You do not have CJKV input methods.
You still have problems with modifier keys becoming stuck "on" because of VT switching races.
* https://unix.stackexchange.com/a/494685/5132
And you also still have ISO/IEC 2022 8-bit character set switching for messing things up. (-:
I was saddened when Spain changed Spanish sorting to make life easier for PC programmers in the 90s. Within only a couple of years that effort became pointless, but was now encoded in law.
”Spanish treated (until 1994) "CH" and "LL" as single letters, giving an ordering of cinco, credo, chispa and lomo, luz, llama. This is not true any more since in 1994 the RAE adopted the more conventional usage, and now LL is collated between LK and LM, and CH between CG and CI. The six characters with diacritics Á, É, Í, Ó, Ú, Ü are treated as the original letters A, E, I, O, U, for example: radio, ráfaga, rana, rápido, rastrillo. The only Spanish-specific collating question is Ñ (eñe) as a different letter collated after N.”
(Copied from https://en.wikipedia.org/wiki/Alphabetical_order#Language-sp...)
While this is true, significant amounts of friction are caused by being unwilling to adapt to the medium's nature. What would writing look like if we had tried to encode sound waves directly onto paper? Instead we adapted our communication to the medium.
* It's 8-bit. (It's 7-bit.)
* It is sufficient for English. (Ð ð Þ þ)
* It is sufficient for Modern English. (zoölogy coöperate £ née resumé)
* It at least has everything that one could type on a typewriter. (½ ¼ ¢ , and that's just starting with some contemporary 20th century U.S. typewriters from IBM such as the Selectric)
* It is sufficient for the U.S., at least. (Not for the U.S. Library of Congress in 1969 it wasn't. It needed 174 characters for catalogue cards, per https://link.springer.com/article/10.1007%2FBF02404378 .)
* It was used by Multics. (Multics used a modified version, padded out with leading zero bits and replacing DEL with PAD. See http://web.mit.edu/saltzer/www/publications/multics/bc-2-01.... .)
* It does not have broken vertical bar. (It was in the 1968 standard at 124, a broken bar because the PL/I language people did not want their unbroken vertical bar to be in the then "national variant" range. https://groups.google.com/d/msg/comp.infosystems.www.authori... http://jkorpela.fi/latin1/ascii-hist.html#7C )
* It does have broken vertical bar. (It was replaced in the 1977 standard by an unbroken vertical bar at 124. Then some subsequent 8-bit character sets in the 1980s re-introduced a broken vertical bar in a second position, with much ensuing hilarity.)
* It did not have arrows. (In the original standard ↑ and ← were where ^ and _ now are, for some consequences of which see https://retrocomputing.stackexchange.com/a/9201/1932 .)
* It is the same as ECMA-6. (ECMA-6:1991 allows national-language variants in several code positions. The standard ASCII characters in those positions are merely the "International Reference Version" and one possibility from amongst several in ECMA-6. See https://www.ecma-international.org/publications/standards/Ec... .)
If you mean, body language is the main or majority channel for communication, well, how much body lanuage matters to communication is contested and differs based on context anyway (see: https://www.psychologytoday.com/us/blog/beyond-words/201109/...).
If you mean, there is a loss of expressibility going from the spoken word to the written word. Is there? You would have to prove that that loss is as significant as the symbolic/representational destruction you are advocating. Much of the imperfect synonyms we have developed in english are used to signal tone and intent -- things that are usually signalled in body language. So, just because we are using the written word does not mean we cannot express those things. You have something in a larger space, you can represent it in a smaller space using longer and more complex series of glyphs.
> Yet here we are still conversing with text despite the technology to record and send each other video having been widely available for a decade or more.
Yes, for the very reason that it erases the extra data. you not being able to see me means that you have to go on my words alone, rather than making extra-contextual judgements about me as a whole based on my physical appearance. That seems only to cement my point, though.
Is there a version of GNU Unifont for Japanese kanji? I could not find any. A few Japanese kanji share the same code-point as Chinese characters despite not being the same.
And this is subjective. I have used Unifont in xterms for years now, and it's very readable and unambiguous to me.
So it's good to see it getting some positive notice.
> It's also a good choice for a dot-matrix LCD/LED display.
Only a high-resolution one.
For example, it's helpful for being able to say "the Arabic name for the Arabic language is spelled ة-ي-ب-ر-ع-ل-ا, alif-lam-ayn-ray-baa-yaa-taa marbuta". But if you want to have an Arabic speaker read it as a word, it ought to be written "العربية" (RTL with ligatures). Your terminal environment is good for the former but not the latter.
The one case where there might be a worthwhile benefit would be for recent emoji additions. But that would be better addressed by a more limited effort to provide an up-to-date emoji-only font, not a resource that attempts (in vain!) to cover the whole of Unicode.
Google's Noto font family is of far higher quality than Unifont and serves the same purpose. Better, too, since there's a lot of details in various languages and scripts that go far beyond "put a glyph here". I've heard only good things about Noto in that respect.
And hasn't been updated since 2017 https://www.google.com/get/noto/updates/
But, other than that...
This is only a useful data point if there have been significant changes to "writing system" released since then.
Multiply that by some 120 writing systems, and you can see why things might get a little big if you want "one font for every language".
(which you can't, the spec doesn't allow for more than 65k glyphs, including virtual compound glyphs, per font. So I'm genuinely confused about the author's claim that they managed to put 137k glyphs in a 65k glyph space. You need multiple fonts tied together using a font family name)
https://github.com/googlei18n/noto-fonts/blob/master/NEWS.md
Every browser and Electron app could bundle it, but this seems like the "wrong" place to implement this. I'd suggest that they could bundle it temporarily until OSes catch up.
Edit to make the comment (hopefully) more valuable: I mean, unless certain choice of fonts is somehow intrinsic to the content or purpose a given website serves, the browser, and by extension the user, probably know better which fonts are preferable for them to comfortably read the textual content.
In my previous app I used web fonts to be able to render Hearthstone cards using the correct font the game uses. In my previous app I used web fonts to
Better still, though, for operating systems to bundle adequate fonts. And indeed, OS-installed font coverage has come a long way in recent years.
Neat idea. I think the transition to UTF-8 is practically done, I'm not seeing � anymore these days (used to be extremely common a while back).
Most systems, when called to display a character which they're unable to render, will render a placeholder. This is most often a dotted box of some sort, roughly the size of a large character. In some systems the dotted box (assuming it's large enough for them to be readable) contains the Unicode codepoint number that the system couldn't render. In a few the box contains some representative symbol that gives you a hint what sort of thing is missing, e.g. maybe it's a Han glyph to suggest that you should look for a Chinese font.
I haven't seen any (they may exist of course) where they render U+FFFD the replacement character �.
The most common reason to see U+FFFD is the reason it was created, something was encoded or decoded in a way that is gibberish and the best option in that case is to replace the minimum chunk of gibberish with U+FFFD and then keep trying. On the Web you'd often see pages which claimed to be UTF-8 but were actually ISO-8859-1 or Windows codepage 1252, neither of which is UTF-8 but they share the most common Latin characters, these days most browsers will auto-detect this goof, and besides most web pages really are UTF-8, but when browsers were less good at guessing and more pages were wrong you'd see it more often.
font-family: Helvetica, Arial, Sans-Serif;
it will fall through. so if a user has Helvetica installed, but Helvetica doesn't provide glyph X, then it will check whether Arial has glyph X.so if you want all non-ascii glyphs to fall through to an alternative font, you need to serve a version of your primary font that only includes ascii.
Here's an example I found a few years ago - https://shkspr.mobi/blog/2015/11/premature-subsetting-of-web... - an English language website never expected their authors to use the é (e-acute) character. So they removed it from their webfont.
[0] or even just a subset if they got "optimised" by anglo-saxon developers who assume ü or ß are useless
* The Standard Unifont TTF Download: unifont-12.0.01.ttf (12 Mbytes)
* Glyphs above the Unicode Basic Multilingual Plane: unifont_upper-12.0.01.ttf (1 Mbyte)
* Unicode ConScript Unicode Registry (CSUR) PUA Glyphs: unifont_csur-12.0.01.ttf (1 Mbyte)
(from http://unifoundry.com/unifont/index.html). And note that unifont_upper only seems to cover plane 1 and plane 14 stuff; they haven't attempted to tackle the plane 2 CJK repertoire.
The article title "Banish the � with Unifont" is also misleading, actually. � is U+FFFD, the REPLACEMENT CHARACTER that typically indicates an encoding error or binary garbage; it's not the same thing as the missing-glyph symbol (often a simple box, though it may vary) that generally appears when font support for a valid character is lacking.
- see that it's foreign text, rather than pictographs or weird english text.
- see when the shapes of other languages are used to make pictures (¯\_(ツ)_/¯)
- if it's a foreign language, you can guess which language
- you can guess at the complexity of what was written, for example by looking for repeated substrings.
Unicode encodes what's necessary for printing books since about 1900 (and a bit more, but that's a fair one-sentence summary). What you want to validate isn't that you'd be able to print every kind of book printed since 1900. You're only interested in some of the alphabets, and you may be interested in more functionality than just printing. For example you may need sorting, or character input with the right sort of interactive appearance changes, or equality testing.
If you decide what you want to work, then googling usually finds a suitable test quickly.
BTW, if you want to discuss which languages's scripts are complex… office, office, office, office.
Many emoji these days are quite complex Unicode sequences with a number of them in the so-called "Astral Plane" meaning they need more than 16-bits to accurately display (proving you aren't treating UTF-8 or UTF-16 as if it was UCS-2), and as sequences include a lot of fun non-visible codepoints ("characters") such as the Zero-Width Joiner, and are very susceptible to breaking if accidentally dropped, reordered, or otherwise spliced (possibly proving you aren't doing back string math or manipulation at the codepoint level rather than the glyph/sequence/combined-character level).
[ETA: Useful sequences to test are any that support the skin-tone and gender modifiers. On Windows, the various "cat occupation" emoji are also interesting sequences such as ninja cat and astro cat. Other platforms have similar unique "fun" sequences that are noticeable at a glance when right/wrong.]
It's not entirely true that if you support emoji well you support any Unicode user's script well, but if you support emoji well you probably don't do anything particularly stupid to make other Unicode users unhappy.
[1] https://www.cl.cam.ac.uk/~mgk25/unicode.html
[2] https://www.cl.cam.ac.uk/~mgk25/ucs/examples/UTF-8-demo.txt