Uppestcase and Lowestcase Letters
tom7.org
tom7.org
Japanese alphabets (kana) doesn't have a concept of upper/lower cases. There are two different types of kanas (round ones and square ones) but they are kind of equal in terms of strength/stress so they can't be used to express anger.
People on 2-Channel forum (Japan's 4chan, basically) came up with a brilliant idea, which is to insert a space between each letter so that it looks a bit wider and has an extra oomph.
Example:
- normal: 今日はいい天気だね。 (it's a fine day, isn't it.)
- angry: 今 日 は い い 天 気 だ ね 。 (IT'S A F*KING NICE DAY ISN'T IT)
Writing word in katakana that usually written in kanji is also works as emphasis with a bit different meaning. It tend to be used as stereotype. For example, 福島(Fukushima) / 広島(Hiroshima) is just a name of prefecture, but sometimes written フクシマ / ヒロシマ that refers nuclear plant accident / nuke bomb event. (I really dislike this usage).
They can be used to convey tones in text similarly to how italics, caps, symbols and other decorations work in general, I think that’s what they meant to say by emphasis.
There are an incredible amount of other ways to add emphasis in Chinese though, so it's not lacking anything, and I imagine the same is true of the other languages.
The way we write modern text is more modern than most people realize, I think. The letter 'j' wasn't really used as a separate letter indicating a separate sound until sometime around the 16th century I think! Case is older, but not ancient. I've seen some Greek on 3rd/4th century middle eastern ruins and I struggle to read the "all upper case with no spaces" writing sometimes, but that's just how it was! No "lower case" until much later...
This was never a goal; look at B ϴ O P Φ Ψ Ω, or on the Latin end C G Q R S.
Inscriptions are very formal; you carve the letters you have no matter what they look like. There was never any difficulty carving curves.
Handwriting is informal; you write whatever you find easiest.
I have a book published in the 1920s with a forward by Stanley Morison, by an author who attempted to enhance the Hebrew alphabet by introducing upper and lower case letterforms to it as well as to bring the letter forms more in line with the styles of the Latin-Greek-Cyrillic alphabets. It's—odd.
Before the eventual standardization of Hangul around early 20th century there were numerous attempts to "linearize" Hangul's characteristic syllabic blocks ("풀어쓰기" [1]). Many of them were influenced by Western alphabets and had two cases, and none were successful. And yes, they are also odd.
https://www.japanesewithanime.com/2018/03/furigana-dots-bout...
However, there is no easy way to enter or display these online, hence hacks like spacing.
Hey Unicode Consortium, where you at?
It would still be a problem of IME and font support.
On input, there's already so many shortcuts and hacks (e.g. SHIFT/CAPS is already taken to force switch from hiragana to katakana) that it's hard to imagine some natural combination that could be memorized.
On the font side, there's already the battle raging for having proper jp fonts in smarphones instead of the chinese ones, so additional support for marginal features is an uphill battle to say the least.
data:text/html,<div style="font:16px monospace; text-emphasis:red filled circle">Hello World
Unfortunately it doesn't seem to be supported in Chrome yet.
data:text/html, <textarea style="font-size: 1.5em; width: 100%; height: 100%; border: none; outline: none; font-family:monospace" autofocus />
mostly useless but I like it
When you land on a site with a delayed paywall, you can copy out the content just before the popup, and then paste it into <body contenteditable> to view at your own leisure.
[1] In this case, Japanese そうですね "I see" written in Hangul.
我!太!满!足!了!
- normal: It's a fine day, isn't it.
- vaporwave: It's a fine day, isn't it.
Not. funny.
this is also used in English but for a mocking tone.
Saying "they didn't have capitals so they used spaces" would sound odd to an alien, who would wonder why capitals were necessary in the first place.
It's fun to hear how it's done in scripts that doesn't support our (as in the average HN reader) default way of doing it (which would be caps or italics/bold).
They contain vowels that go beyond the basic 5 "aeiou" of the English language (eg. see [0]). These suffice to let the speaker know how exactly to say the word unlike in English where it has to be learned case by case based on whatever is popular or acceptable pronunciation.
A subset of of these languages are tonal languages which also have special characters to additionally allow the speaker to set the pitch of the word correctly which then changes the meaning of the words [1].
[0] http://www.bbc.co.uk/languages/other/hindi/guide/alphabet.sh... [1] https://en.wikipedia.org/wiki/Tone_(linguistics)
Even English itself has vowels that go beyond the basic "aeiou" of "the English language".
* the letter Y
* vowel pairs, like ‘ou’
* accents, like in ‘précis’
Is that what you mean?
But, you say, that's 20 lexical sets, not 13–15? Well, no dialect distinguishes all 20. My idiolect (a slight variant of General American) realizes TRAP and BATH as [æ], PALM and LOT as [a], CLOTH and THOUGHT as [ɔ], KIT as [ɪ], DRESS as [ɛ], STRUT as [ʌ], FOOT as [ʊ], FACE as [ei], GOAT as [ʌu], FLEECE and HAPPY as [i], GOOSE as [u], PRICE as [ai], CHOICE as [ɔi], MOUTH as [æu], and COMMA as [ə]. That's 15, or 12 if you leave out PRICE, CHOICE, and MOUTH, which are diphthongs made of vowels that also occur isolated. (GOAT is debatable, usually analyzed as [oʊ] or [ou].)
Different dialects draw the boundaries in different places; for example, dialects with the "trap–bath split", such as RP, famously realize TRAP and BATH differently ([æ] and [a] in RP). Some dialects have fewer vowels; if we consider Indian English to be a single dialect, it may have more speakers than even GA, and most varieties of Indian English have fewer vowels than 12. I haven't found a good phonological analysis, but if you know any Indian English speakers and also know phonology, you know what I mean. https://en.wikipedia.org/wiki/Regional_differences_and_diale... goes into some detail.
______
† The historical gap might be much larger than this. Sumerian cuneiform and Egyptian hieroglyphs date back about 5300 years, and they provide evidence that spoken language was considered to be universal among humans at the time—there is no suggestion of tribes that lacked language anywhere in the written record. Today there are still peoples without written language, and a few who only acquired written language within the last generation. So we have good evidence that it has taken at least 5300 years. But Homo sapiens has been around for sixty times that long, over 300 millennia, and stone tools date back 2 million years. It strains credibility to imagine that the authors of the Lascaux cave paintings or the Denisovans who invented sewing were so unlike us as to lack speech; the origin of spoken language is usually dated to before 40kya. Unfortunately, no tape recorders have yet been found from that epoch, so the uncertainty of the antiquity of spoken language ranges over nearly a factor of 100. Maybe spoken language is a million years older than written language, or five million. Probably not ten million, though, or we'd be studying chimpanzee folklore.
This is more due to how widespread English is, and how vowels/pronunciation have shifted over time. For example, The Great Vowel Shift. [0]
Korean has exactly the same issues, albeit to a smaller degree. There are plenty of words that aren't pronounced like how they're spelled, due to grammatical rules. 종로 as an example. Or cases where words sound exactly the same and you just need to know the context/spelling- 쫗다, 쫒다, 쫃다, 쫏다, 쫓다.
Then there's regional slang/pronunciation/dialect. Busan dialect is fairly different from 표준어, or the "official" standard language. This phenomenon is not unique to English in any way. Any language scaled up will develop these issues over time.
At least you can try to pronounce "cough" or "종로", instead of not being able to pronounce "願" at all because you don't already know the pronunciation.
I'd go so far as to say that english is now spoken by so many people in so many different regions, all of whom can now be heard by each other on a reasonably frequent basis, that it has forced english speakers to become so adept at vowel-reconstruction that one could pronounce words with completely arbitrary vowels and still be understood.
This was still occurring before people were able to widely hear other regions' speakers, though.
However things like cough, plough, although, thorough, etc having different, -correct- pronunciations are due to English taking in words from other languages.
You can have three things:
1. A spoken language that evolves over time.
2. A writing system that accurately describes pronunciation.
3. A writing system that indicates history and etymology.
But you only get to pick two. English went with 1 and 3, which is arguably the optimal choice.
1. Because it's more important to know what words mean than it is to know how they sound.
My daughter reads a ton and learns a lot of words from reading. Fairly often, she mispronounces them, and that's OK. What's more valuable is that she can often infer the correct meaning of the word both from the surrounding context and from the parts that the word is made of. If we normalize spelling to match pronunciation, much of the latter gets lost.
It's easier to see that "mean" and "meant" are related than "meen" and "ment". "History" and "story" versus "histery" and "story".
2. Because pronunciation changes over time. If we continuously change spelling to match, it means older printed works get harder to read. In the worst case, they can appear to be saying different words than they intended.
3. Because pronunciation isn't uniform across regions.
Should "lawyer" be spelled "loyer" or "lawyer"? Is "crayon" spelled "crayon", "crayawn", "cran", or "crown"? Is it "caramel" or "carmel"?
> Japanese alphabets (kana) doesn't have a concept of upper/lower cases [...] so they can't be used to express anger.
It reads to me a bit like "the natural way to expess anger is uppercase, but they didn't have that, so they did the other thing instead".
Would you find a sentence like "English didn't have dots so they used uppercase to express anger" equally natural?
The comment is premised by how in japanese there are "two" sets of "letters" with slightly different uses but differently from lowercase/uppercase the difference does not translate in emphasis.
Uppercase letters are linguistically related to emphasis independently form net-speak (proper names, I, beginning of sentences), so to me it reads as "the thing that naturally works here is impossible there, so they have a different solution to the problem".
It is fine to take the perspective of your own cultural situation.
Maybe it seems so odd to me because there are so many licenses, contracts, or government forms that use ALL CAPS as emphasis. I don't know.
It just doesn't translate to shouty in reading mental voice.
To me (and probably most others), license texts and such absolutely look like they are shouting. I do not understand whence the convention of having them in ALL CAPS, and can only assume it's in itself some sort of a cultural association between ALL CAPS and IMPORTANCE. I have only seen it in English legal texts, anyway – is it even used in other languages? It looks to me like ALL CAPS was originally used to emphasize key points, and then an inevitable race to the bottom happened until EVERYTHING WAS IMPORTANT which really means that nothing is important.
When it comes to typography, nearly every type of emphasis employed in Western text except italics (and sᴍᴀʟʟ ᴄᴀᴘs, which see too little use these days methinks) only exist due to technological limitations, particularly the extremely limited typographic options available to typewriters and, later, 7- or 8-bit text terminals. This includes ALL CAPS, s p a c i n g, and u͟n͟d͟e͟r͟l͟i͟n͟e͟d͟, never mind ASCII crutches like /pseudoitalics/, _pseudounderlined_, and ∗pseudoboldface∗.
Honestly I'm a little surprised that people who've likely seen their share of BAD COMMAND OR FILE NAME would still read that as shouting.
Is there a place where I could learn more things like this one?
กาก -> fail
ก า ก ก ก ก ก -> FAIL!!!!!11!
Tom, if you are reading this, I am curious what a capital smiley face emoji looks like. Or the capitalest. What does nirvana look like in a twittable small yellow circle? Or a lowercase one. Or lowestcase one. What is the graphic depiction of the pit of despair.
You got me thinking. Sure, emojis are in color. But could you make them black and white? Or apply the transform to each color channel independently. I'm convinced there is a way, and if there is anyone who can find it, it is you.
And if you do go this way, why stop at emojis? The world needs to know what a capital Mona Lisa looks like. Or the capitalest.
I think you're on the path to general artificial intelligence here. iA.
Also, check out the author's video: https://youtu.be/HLRdruqQfRk
For more non-serious research that is often executed seriously, see: http://sigbovik.org
I do wonder whether the problem is that training on multiple fonts just doesn’t actually add more data points to the dataset. Fundamentally, you’re going to get a model that knows how to uppercase or lowercase each canonical Latin letter. This model might actually be quite good at generating appropriate lowercase forms for a font given the uppercase glyphs, but it’s not learning an abstract concept of ‘lowercaseness’.
Instead of extending the training set with more fonts, the only real source of additional training data would be more alphabets. Cyrillic and Greek both have case systems so just adding those would have more than doubled the number of training cases - I think the trick would be to start from the Unicode case mapping tables to generate your example data, to give your model more variety of upper/lowercase pairs to get its teeth into and really make it possible to ask it to uppercase an arbitrary letter form.
If you asked someone else, "I want to automatically turn characters into 'more uppercase' or 'more lowercase' versions, and it has to involve neural networks; what do?", they would say something like "Easy enough! dump character maps into a standard StyleGAN2-ADA, train for a few days, then find the latent direction corresponding to uppercasing & lowercasing, and edit whatever font you please." (People have been generating fonts with RNNs or GANs for ages.) You could do this over a week or so, the tooling has gotten pretty easy to use.
And if you asked someone on the cutting-edge, they'd suggesting using OA CLIP through Aleph/BigSleep/etc to automatically edit images using a text input prompt of "a lowercase letter" or "an uppercase letter", or one of the hybrids like StyleCLIP https://github.com/orpatashnik/StyleCLIP . This approach might take all of an hour or two (but results probably would be worse).
It would be interesting to see if AI could be used to fill in some of the gaps in the fonts - ie: numbers for his Angstrom font (https://www.dafont.com/angstrom.font)
Someone could also train a model that given some glyphs of a font predicts the rest of the glyphs. Then we can do weird things like give it glyphs from multiple fonts as input to make a hybrid font.
Do you mean relationship in terms of appearance or pronunciation?
Я maps to "ya" sound in Russian, И is "ee" and Н is "N" in Russian and "ee" in modern Greek but in ancient Greek it's something like "e" in "bed" in American English.
Some fonts do not have Cyrillic glyphs. If they could be generated by an adequately trained model based on the other glyphs, then that font's multilingual coverage could expand automatically.
Do give this a chance
EDIT: I forgot what day it is. I really like the idea though and I think with a little more work it could really deliver
The database is just filled with garbage that is unusable for this project: Fonts that are completely illegible, fonts that are missing most of their characters, fonts with millions of control points, Comic Sans MS, fonts where every glyph is a drawing of a train, fonts where everything is fine except that just the lowercase r has a width of MAX INT, and so on.
This guy is the Terry Pratchett of home-grown deep learning libraries...
So I did that and let it run for a month. Actually I had to start over several times with different parameters and initialization weights because it would get stuck (Figure 11) right away or as soon as I looked away from the computer. I prayed to the dark wizard of hyperparameter tuning until he smiled upon my initial conditions, knowing that some- where he was adding another tick-mark next to my name in a tidy but ultimately terrifying Moleskine notebook that he bought on a whim in the Norman Y. Mineta San Jose International Airport on a business trip, and still feels was overpriced for what it is.
https://drops.dagstuhl.de/portals/lipics/index.php?semnr=160...
This is not as absurd as it sounds. After all, over the centuries, we have introduced punctuation to make it easier to read, unlike classical writing [1]. Why not carry through with remaining reforms?
Like how some advertising material renders "$499.99" as 499 in a large size, followed by 99 in a smaller, often superscript style, and no discrete decimal.
Capital letters in modern script are therefore not a remnant of technical limitations, the mix of lowercase/uppercase was actually voluntarily introduced for legibility once these technical limitations were gone.
Funny thing to understand; this has no connection to reality. Letters for inscription are formal and there is no bias towards angularity. Look at this plaque from the first century: https://static.timesofisrael.com/www/uploads/2012/04/Roman-m...
You can clearly see the letters B, C, D, G, O, P, Q, R, and S graven in bronze with all the same curviness they'd have if they were instead carved in marble. [1] They're carved that way because that is what the letters look like; if they're difficult to carve, that just means the carver has to suck it up.
The letter we commonly render U is V in Latin. That is not due to the needs of the medium; that is what the letter looks like. In Latin, there is no letter U. V is not pointed because it was difficult to carve C, D, O, Q, P, R, and S. V is pointed for the same reason M is -- because that is the shape of a V.
[1] Here's an example that is carved in marble: http://codex99.com/typography/images/ancient/trajan_sm.jpg
At any rate it's not at all obvious to me that removing case would improve things. I'm currently reading a book my Valter Hugo Mãe who's an author who (generally) only uses lowercase letters: https://svkt.org/~simias/up/20210402-153754_maquina.jpeg
I've also got a Latin grammar book that, for extra authenticity, only uses uppercase at the beginning of the paragraphs: https://svkt.org/~simias/up/20210402-154010_latin.jpeg
In both cases I often find myself missing the end of a sentence, I think mainly because ',' and '.' are easy to mix up but normally you expect '.' to be followed by an uppercase. Of course part of the issue might just be lack of familiarity.
I'm not arguing that capitals are vital and we couldn't read and write correctly without them, but as you point out we could say the same of spaces and general punctuation. I can read unaccented French just fine for example, but it does force me to slow down at times and to infer more from context.
butanywaythatsjustmyopinion
As applying cases is then just 'addition', then you can likely get the 'multiplication' and 'divisions' of the cases as applied. 'Exponentials' are just around the corner too.
Is that is the situation, that we can now get transcendental cases. The 'pi' case, or the 'e' case, pick your favorite transcendental.
More interestingly, you can then pull out the imaginary case, using Euler. e^(i*pi) = -1. Maybe that the lower-er case is just the 'e' case to the power of the 'pi' case times the 'i' case. Whatever those operations may mean.
Of course, one you're at the imaginary cases, you might as well step up into derivatives and integrals, it's just curiosity after all. Then you'll be doing partial derivatives and then Lagrangians.
Eigencase-ing comes next, and then Maxwell's equations in case format. 'Del'ing your cases should be a real trip.
After only a bit of puzzling, you're doing quantum mechanics operations with your cases, because why not?
Case-ing operations have a fruitful future for any mathematician, it seems.
But he's a great coder and had loss of successful projects. It's just the way he learned and it doesn't seem limiting or annoying to him.
Now though, I have the opposite bad habit, if I need to write a long string of uppercase letters, I don't turn on caps-lock, I instead keep my left little finger pressed on the shift key.
... fonts where everything is fine except that just the lowercase r has a width of MAX INT...
Okay, I cannot imagine what the use of that font would be, but I really want to know.
My only suggestion for improvement for the modeling is in the part where you tried to find the "ideal" letters for the model.
Instead of sampling generated random bitmaps and ranking those, you could instead initialize the inputs to random continuous values, then optimize the output score for the desired letter with respect to the random input (with your fixed model).
Indeed, this is how people get those trippy dog images you've probably seen.
Anyway, this video was just so good. I'm amazed.
[1] https://www.youtube.com/channel/UCS0N5baNlQWJCUrhCEo8WlA
- William Zinsser
Zinsser was actually talking about writers here. And it might be a bit hyperbolic, sure. But I think the fact is that people who love to program spend a lot of time staring at words, and given a chance they'll take interest in the clothes that words wear.
It's just that I'd be interested in the main fonts that make practical sense for projects and not so much on the quest for the holy grail of fonts that makes all men fall to their knees in awe at its aesthetic perfection. Doesn't matter how pretty you write bad content, anyways.
> Probably someone already had this idea and did it before I was even born