Why are letters shaped the way they are?
vice.com
vice.com
It wasn't until the digital revolution that the 7-segment pattern belying most letters was apparent. There is no "7 segment font" for the Guttenburg printing press. Most of people through most of history would not have consciously known there were atoms of writing smaller than a single letter.
Vocabulary almost certainly has evolved in a way that efficiently compresses semantic information. We just haven't noticed the 7 segments yet.
And the accented letters in Spanish just denote a stress mark when the spelling is not following the rules.
Cortes is "CORtes", it means "courts".
Cortés is "corTES", it means "courteous".
Then, in ASCII, it could be represented as Cort'e's or cort|e|s or any other small symbol.
The diaeresis over ü would be a bigger task, but not impossible.
7 segment font has been designed to work with 7 segment display designed to accommodate a simplified-lower-density English alphabet.
It's a typical engineering feat that says not much of the source alphabets/calligraphies: if you were used only to cursive, kurrent or gothic calligraphies, 7 segment is not much relevant.
If you are using greek alphabet, you also get into trouble to differentiate between latin and greek characters.
Not speaking of other languages representations.
>It's a typical engineering feat that says not much of the source alphabets/calligraphies
It says everything about alphabets/calligraphies. We could have lived in the world where each of the 26 glyphs was as unique as a frame of TV static, and there was no way to compress them down into segments or anything else. We didn't end up in that world. We also didn't end up in the world where clever engineering got each letter down to 12 segments, which is fine but still has twice as many bits in the glyph as there is information to encode. No, we ended up in the world where the "typical feat of engineering" was possible, and thus have a representation that only wastes 1.36 bits per letter.
An engineer could not have willed a system of k-segments into existence and imposed it on users. An engineer cannot make things look like letters if they do not already look like letters. The 7 segments had to be the subconsciously recognized abstract form of letters this entire time.
>if you were used only to cursive, kurrent or gothic calligraphies
then there would be a different subconscious abstract form of the alphabet and a different near-optimal k-segment display would have been invented.
They did reduce something that existed to a lower resolution so that it fits into k-segments: how do make o (lower case o), O (upper case o), 0 (zero), D (upper case d), Q (upper case q) coexist on a 7-segment display without ambiguity? You start by deciding that lower/upper case it not important and that only one case for each letter will subsist. That's reduction.
> The 7 segments had to be the subconsciously recognized abstract form of letters this entire time.
There are other languages of which the scripts do not fit in this view. English Roman-derivative alphabet is not "better" because it compresses itself more into a specific-imposed compression format.
https://gankra.github.io/blah/text-hates-you/ from https://news.ycombinator.com/item?id=30330144 is a perfect illustration.
Actually, that example is perfectly doable on 7 segments.
>coexist on a 7-segment display without ambiguity?
BuT WhAT AbOuT UPpEr AnD LoWEr CaSe d aND q??
No information was lost. It doesn't matter that some glyphs didn't make the cut, because those glyphs were redundant. The engineer didn't get to choose which glyphs they could safely ignore without destroying the underlying message content. That's something that is intrinsic to the underlying alphabet.
Also notice I only included one copy of the 26 letters in the English alphabet bitrate. If you insist that We SHouLD hAvE CApItAl ANd LoWeR CAsE LeTTeRs then the symbol count balloons to 70. That's 6.13 bits per letter which only leaves .87 bits of code space left in a 7-seg system. It should come as no surprise the capitals won't fit. It would be miraculous if they did! I'm sure the precious CaPItAl LeTtERs can be saved by adding an extra segment or two.
>There are other languages of which the scripts do not fit in this view.
Are there, or has no one thought hard enough to figure out the k-segments that would work for those scripts.
>English Roman-derivative alphabet is not "better"
Stop putting words in my mouth. I didn't say anything about anything being better. I said that there is evolutionary pressure towards glyphs simplifying down to the minimum amount of features needed to tell them apart. There is nothing in this hypothesis that is specific to English or even specific to the 7-segment display. This is a purely information theoretic explanation for why the shape of letters evolves in such a way that eventually you could invent something like the 7 segment system.
Diacritics are the obvious counter example. The two dots you can put above a, o, u to change their sound, or in other languages above any vowel to indicate it's part of a different syllable then the following vowel. The undercomma to slightly change the sound of a s or c. The accent grave, acute or circumflex to indicate if that e is to be pronounced é, è, or ê (so more nasal or open). And that's just the popular, European ones. Languages like Arabic make diacritics even more central to their writing system.
IMO, the various letterforms in use by all of the alphabetic languages have more to do with the available writing technology of their day, and in keeping the letterforms sufficiently distinctive to quickly recognize whole words. Compare Etruscan engravings to early Roman engravings to German blackletter to Spencerian cursive to Zaner-Bloser cursive for example.
Hangul has a similar diversity in stroke shapes to Latin characters, but much more diversity in stroke position. I pity the poor engineer who might have been tasked with constructing an N-segment display for Hangul in the days before dot-matrix technology was feasible.
Finally, the 7-segment display doesn't actually look particularly recognizable when being used to render latin text. It only really works for the European evolutionary branch of Arabic/Hindu numerals. Even with only 26 (modern! English!) letters, we still have enough stroke diversity that we'd prefer to render text with more complexity. 2^7 segment combinations for only 2^3.3 characters doesn't sound so optimal any more, does it?
But if you want to stick with your information-theoretic observation, I'll point out the completely unrelated fact that some of the early and successful forward error correction schemes for Gaussian channels use rate 1/2 coding. Anything optimized for burst errors is quite different, though.
I agree that we don't need upper and lower case. In fact you may notice I did not include the 2 cases in my calculation of number of glyphs in English. I contend that of the 26 modern letters, at least 20 (and up to 23) can fit on the 7 seg display and still be recognizable. Granted the 7 seg alphabet would be forced to use the (currently) upper case form for some letters (like A), and the lower case form for others (like q), but at least one of the two forms fits the 7 segments for every letter except M,T,K,W,V,X. Since we already accept not having two forms for each letter, we may as well re-assign the former capital U as a "double sized u" aka W. As M is W upside down, the same trick applies to make a 7-seg M. Yes I've mangled the letters, but the outline of the shape is close enough that you won't trip up reading it in the context of a word. There's already some precedent for using "7" to replace T (as in 1337 5p34k). So that just leaves K,V,X as "not recognizable on 7 segment display".
But you've kind of made my point for me.
> and in keeping the letterforms sufficiently distinctive to quickly recognize whole words.
Just by evolution the letter forms evolve towards the minimal amount of distinguishable features needed to encode the whole alphabet. Evolution is an optimizer not a designer. The fact that there are still 3-5 letters left that didn't get optimized into 7 segments is what we should expect from such a process.
Who uses only letters? You want an alphanumeric display, to show 26 (27,8,9...) + 10 = 36 (37,8,9...) characters. That's your actual everyday "alphabet".
And then you're totally screwed from the start: 0OD, 1iIl, 2Zz, 5Ss, 6bG, 8B, 9qg...
My claim is that the 7-segment display is only usable for its original design scope: the numerals 0-9.
IPA? A perfect dictionary would have a one-to-one mapping between letters and sounds. English is both redundant (c-k, c-s) and lacks letters for certain sounds ("ch").
Your choice of punctuation, most of which cannot be represented on a 7 segment display, is arbitrary[1], ASCII/US keyboard centred, and only added to your argument to increase bit count.
[1] why include /-+ but no *, =? # but no @, &, $, %? () but no []{}<>? " but no '?
Why /-+
Don't interpret them as math. Those are "or" (as in yes/no), hyphen, and "and" (as in you+me). I rejected ampersand because the use case was covered by +. I included one set of parenthesis because parenthesis are a gramatical necessity as scope can otherwise be ambiguous, but more than one style of parenthesis is redundant. I reluctantly added # because in order to fit everything in 7 sev, you are sometimes forced to use 1337 5p34k. So it might be necessary to have a symbol for "read this as a number". Contrary to your assertion that I was just padding out the count, I was actually trying to do the opposite and keep the extra symbol count down to only what is needed. You could cut all the special characters and still be at over 5 bits.
You said T and V twice. I already admitted elsewhere V,X,K are problematic. G -> g -> 9. Z -> 2. S being indistinguishable from 5 means you are forced to accept collision between numbers and letters anyway, so by the conventions established in 1337 5p34k, T -> 7 is an acceptable substitute. Accepting that there is only one form of a letter, no more capitals, you can do "(" with what used to be C, and ")" with a backwards "C". "/" is absolutely doable with 7 segments.
Of the 50 symbols I started with, there are just 9 that don't really fit in a visually intuitive way. KVX.,;+#
And honestly that's fine. The point of this exercise wasn't alphabet reform. The point was that 7 is a close estimate for the number of bits needed to define the shape of a letter in the English alphabet. The fact that the 7-segment system can cover just over 80% of these glyphs without anyone intentionally designing it that way is suggestive of the idea that alphabets naturally optimize towards some system of n-segments.
>You cherry pick between lower and upper case letters, discarding half of them, only to shoehorn them into your model.
Not really sure whats the problem here. I've been consistent. 1. The distinction between upper and lower will no longer exist. There is just one glyph per letter, sometimes the glyph formerly known as "upper", sometime the glyph formerly known as "lower". 2. Some glyphs that might have been used as an upper/lower version are instead used to encode something that otherwise wouldn't have a glyph (like m, w, parenthesis).
It's not shoehorning to repurpose the symbols I previously removed for redundancy.
>The way you proposed to represent M and W is non-intuitive
Try writing out a few sentences this way. I assure you it has enough in common with the outline of M and W that you won't trip over reading it.
Yes this is a bit of a stretch but when you write it out they look similar enough to the correct shape that you can still read the words just fine. The letters that don't really fit in any way that looks right are x,v and k.
The short answer is that it's a mistake to read my comment as saying 7 segments per se in the usual figure-8 shape is universal to how brains comprehend writing. It's not. You can dig up counter examples all day because that wasn't my point.
My point was that for a given language, the writing system tends to evolve to the point where each glyph contains just enough distinguishing features to not confuse it with any other glyph. There is also selective pressure to reuse the same segments but in different combinations, rather than try and come up with wholly different lines for each glyph which can only be learned by rote memorization. The end result of these pressures is that any language will evolve towards its own version of a 7 segment display. Call it a k-segment display, where k is only slightly higher than log_2(the size of the alphabet). I don't know what the k segment display looks like for Cyrillic or Greek or whichever other alphabet you pulled the other glyphs from, but I'm sure one could be found.
It definitely exists for Korean, as the Korean alphabet was designed in the 19th century and the sub-letters are intentional. I mention this because one of your chosen examples is a Korean letter.
This is absolutely not true. If this happened, you would have a very bad writing system.
Human communication systems have tons of redundancy in. The actual regular alphabet used by English has tons more shapes and detail that you are accounting for, and it uses it incredibly inefficiently. You can create a huge number of additional letters using the shapes that the alphabet is using.
And this is a good thing. It adds redundancy, and it very specifically is not the case that "each glyp contains just enough distinguishing features to not confuse it with any other glyph" - in fact, it contains far more distinguishing features than you need! This is to allow you to recover the information when things are unclear. There is excess information encoded as a form of error correction.
This is true for writing systems, for language grammars, for vocabularies, for every part of human communication. We always over-specify, to allow for understanding even when the receiver misses out parts of the information.
1 bit of error correction for 5.6 bits of information sounds pretty good.
False, it was in 1443. https://www.korean.go.kr/hangeul/setting/002.html
Clearly not every letter in the canonical English script can be rectilinearized onto the 7 segment display even if the alphabet can be encoded in 7 bits. That is a significant difference between being able to encode it and being able to represent the existing script.
You have to admit, almost all of them can be. Including the numbers. There are a few exceptions of course.
Most kids today who learn to write learn to write the Latin alphabet learn by tracing individual strokes or “atoms” as you put them. Minimizing strokes introduces cursive.
Stepping back into history, there’s a lot of overlap between how you would carve letters into a clay tablet/stone and how a seven segment display works. So that’s a fun parallel.
I learned a similar concept in Typography class... for an assignment to design a new number. I was a little indecisive and apologized when I showed up to critique with two different number drawings [0]. My professor (Cyrus Highsmith [1]) looked at them and said not to worry because they were the same number anyways. I didn't get it at first, but a few minutes after sitting back down I realized what he meant, and it's stuck with me ever since.
Using the 7-segment display as an example would have been a great way to explain the concept to me at the time. I'll remember that for next time!
Also a cool way to think about it, because it begs the question of which 7-segment permutations don't map to existing symbols, for whenever someone needs to design a new one in the future.
Applying these constraints, There are a handful of could-be letters left over. Some are already in use like the upside down A (used in mathematics), Capital Gamma, entailment symbol (used in logic). Others, like upside down and/or backwards F, upside down 4, and -| aren't used anywhere AFAIK.
upside down A: ∀
Capital Gamma: Γ
Entailment symbol: ⊧
That last one would also be hard to represent in a 7 segment display, perhaps that would be the upside down F
> We have a 26 letter alphabet
Who are "we" here?If someone says "Most..." or "Nearly all...", they already know there are exceptions.
Everyday writing in Ancient Egypt used the Heiratic script, which then evolved into the Demotic script, and then finally the Coptic script, which is the Egyptian language written with Greek letters (plus 6 extra letters for sounds the Greek alphabet doesn't cover).
The Coptic language is no longer the official spoken language in Egypt (replaced by Arabic), but it's still used today liturgically in the Coptic Orthodox Church. Sort of like how Latin is still used in the Catholic church.
What's particularly interesting is that the Greek alphabet was based on the Phoenician alphabet, which was in turn derived from Egyptian hieroglyphs. Full circle!
Also this misses the point. The division of the letters into 7 segments or into some other system of segments or dots or something else is immaterial. The point is that a writing system will converge towards having as many distinguishable features per glyph as there are bits per character. Any curves or embellishment to a glyph that does not add more information can and will be ignored and forgotten. A glyph that reuses elements of other glyph's will be easier to learn and thus more likely to be adopted. The selective pressure on a writing system is to converge on something similar to the 7 segment display.
The problem is that you consider the current state of one specific language as an optima.
> Any curves or embellishment to a glyph that does not add more information can and will be ignored and forgotten.
Well... you might not have been of the generation that did write a lot by hand. There's a whole lot amount of information you can embed into the way you write/draw your letters, even in the density/dryness variation of ink and the variation of size, inclination, padding, etc. 7-segment display are an accident between manual written text only and today's high density displays.
It's reductionist such as saying a zoom call is definitely an optima to replace physical meetings. While this is true, it highly depends on the nature and content of said meetings. The raw difference in bandwidth and quality of information is significantly huge.
> The selective pressure on a writing system is to converge on something similar to the 7 segment display.
And what if we had a different system that 7-segment display? (which we have, by the way, so better take advantage of that)
I consider it near to an optima for encoding English in particular. A different spoken language, which might carry a different amount of bits per symbol, would be best encoded by a different writing system. I'm sure Cyrillic is near optimal for encoding Russian. I don't know enough to tell you how to break Cyrillic into k segments and still have it be recognizable, but it can probably be done.
>There's a whole lot amount of information you can embed into the way you write/draw your letters, even in the density/dryness variation of ink and the variation of size, inclination, padding, etc.
This is one of those claims like "most communication is non-verbal". But really, what's the bit rate of information encoded in writing style? You say "a whole lot amount of information", but how many bits is a whole lot? And in any case, how much of that subtle, undocumented writing feature would one expect to find in a given document throughout history? Were the monks and scribes and masons who were commissioned to inscribe all the civic and religious texts also inscribing all that subtext, or was their handwriting style the result of trying to write as quickly and efficiently as they could manage? Obviously it was the latter.
To say the 7 segment display is an accident is to miss the point. No matter what engineers might think is desirable, there was no a priori reason to think that letters with all their different curves and forms would be compressible. We could have ended up in the reality where there was no way to display the letters of the alphabet other than pixel by pixel at high resolution. It could have been that each letter was as unique as a QR code, and anything less than transcription of each black and white dot would fail to suffice. But we didn't end up in that reality. We ended up in the reality where all the information needed to discern the character compresses down to 7 bits, and that's just slightly more than the 5.64 bits represented by the character itself.
>It's reductionist such as saying
Well call me a reductionist then.
>And what if we had a different system that 7-segment display?
I couldn't agree more. Flag semaphores make use of two flags held at up to 8 different angles (up, down, left, right, and all 4 diagonals). That gives a code space of 8 choose 2 which is 28. Adapting a similar system for an alphabet would be ideal. Every letter is just two strokes, one coming from the edge into the center point followed by another from that center point to a different edge. Alphabetical order would be easy to remember since it would work exactly the same as hands ticking on a clock. The alphabet could easily expand with 3 or 4 handed characters if needed. The positions of the some of the segments could encode information such as "this is a vowel" or "this is a grammatical particle" making the system that much easier to learn. The system would obviously lend itself to sign language making it more practical to communicate with deaf people.
I'm getting confused. At first you made this about written alphabets, but now you somehow relate this to information in the spoken language. The thing is, spoken language contains significantly more information than 26 letters. There are a number of sounds which are associated with combinations of letters e.g. "th". I have to agree with what someone else said further down the thread, this seems to be much like numerology after the fact that someone engineered 7 segment displays to represent ASCII characters.
Over a whole copy, high enough to distinguish the handwriting or writing style of different persons, and to probe whether a text is a first account text, a contemporaneous copy or a forged later one?
> And in any case, how much of that subtle, undocumented writing feature would one expect to find in a given document throughout history?
There's a reason original artefacts are kept with great care and their high resolution copy have more interest than the full text-only transcription (which itself can be subject to interpretation in some cases).
> Obviously it was the latter.
Was it? Were ornamented, illuminating manuscripts about quick and efficient production? or about effective and subtle communication?
Particularly in the case of illuminated scripts, where illustrations could bring either additional context or critique of the first-order reading of the text, either visual support to the text meaning to the less literate.
I add the caveat that traditional vs simplified Chinese suggests I might be wrong on this specific example, and that institutional stability and standardization froze Chinese writing in place before it naturally evolved to a simpler state.
It's true that visually speaking, some sound parts consist of for example boxes or lines that are also used in other characters, but they have no independent sound, or even deeper meaning.
"Hello" in ...
ನಮಸ್ಕಾರ (Kannada) नमस्ते (Hindi) வணக்கம் (Tamil) ہیلو (Urdu) こんにちは (Japanese) ନମସ୍କାର (Odia)
One of the compromises you have to make to get 7 segment writing system is that you won't always have capital and lower case options. Rather, some letters can only be written lower case, some only upper case, some both, and some sort of both but one of the cases conflicts with a different glyph (eg lowercase l and i). Eliminating case entirely and just having one symbol per letter is fine since capitalization isn't encoding much anyway. Another compromise/hack you can do is accept 1337 5p3ak as de jure spelling. So the number and letter characters are literally the same, and reading them as one or the other depends on context (or perhaps a special "this is a number" glyph like "#").
I admit some of the transpositions into a pure 7 seg alphabet feel like shoes that are slightly too tight. But we can break them in with time!
This is not a mystery, and since the modern letters ultimately got their shapes from pictographic drawings, there is no real connection between shape and sound anymore. At most it may have influenced how the letters have been simplified over time.
It even looks like a foot, which is fun. I took my daughter to the Met this weekend looked at some Egyptian artifacts and was watching some videos about hieroglyphs after, so I don’t actually know much about them, but they happened to have mentioned this on one of the lectures.
Remember, most Hieroglyphs had two or three consonants, and there was no Ancient Egyptian "alphabet", that is something they made up for tattoo's and tourists. When you wrote something in Egyptian, you didn't just combine a bunch of characters that each had a single consonant sound, that was an innovation of the semites. Then later the Greeks come up with the concept of vowels.
Apart from being shaped vaguely like a house, another clue is in the name. Semitic B's have a name that means house.
I don't know too much about this, so I definitely could be mistaken, but the professor seems to say this hieroglyphic may be connected with our current letter "B". Also, I think he said something about "house" on another video and I believe it was different.
When it comes to that video, surely the professor would be aware that lowercase letters are a late invention, b is just a version of the original B, so the lowercase shape per se is irrelevant. Furthermore, I don't think there is any doubt in anyone's mind that B comes from semitic beth, so whatever you do, you need to compare THAT shape to Hieroglyphs.
𐤁 in Phoenician, ב in Hebrew.
Check out https://en.wikipedia.org/wiki/Bet_(letter) and scroll down to the comparison between Egyptian and the Semitic languages.
And again, there's also the name of the letter, which means house, not leg.
The Wikipedia article for "b" seems to list both the "foot" and the "house" hieroglyphs in the development on the sidebar... not that Wikipedia is always correct.
I think the professor in that video is quite well known and has written several books on Egyptology, so I doubt he would be completely off base on something so basic. In any event, you're absolutely right that there the house hieroglyphs is related. I was wrong to question that. But it seems that the "foot" hieroglyph maybe has some relation too. If you find anything more about this, I'd be curious to hear what it is.
"The Egyptian hieroglyph for the consonant /b/ had been an image of a foot and calf ⟨ Leg glyph ⟩, but bēt (Phoenician for "house") was a modified form of a Proto-Sinaitic glyph ⟨ Bet ⟩ probably adapted from the separate hieroglyph Pr ⟨ Per ⟩ meaning "house"."
It doesn't seem from this that they are claiming that B comes from the leg glyph. so it's very strange they put it in the side-bar, if it's etymologically unrelated.
Also, if you look at the Aleph article, https://en.wikipedia.org/wiki/Aleph, they have a sub-section with some Hieroglyphs that are pronounced as aleph, without mentioning any historical connection, which seems very strange. The rest of the article talks about letters that are historically related, so why bring up the vulture glyph? In the article for a they don't show the vulture though, only the ox.
I think part of the reason people say things like this is that they are so strongly inclined to see writing systems as alphabets, so they just have to find "the letter b in Egyptian", even though that makes no sense at all. Egyptian didn't have a letter b, they had a number of glyphs that represented 1, 2 or 3 consonants.
Btw I'm sure the professor knows his Egyptology, so the preceding wasn't referring to him, but maybe he didn't study the origin of the Semitic writing systems. That's not really Egyptology.
It's a bit like doing this:
"It was a very sunny day." <- regular sentence
"It was a very [sun emoji] day." <- same thing
"It was a very sunny day. [happy face emoji]" <- v3 pictos that (implicitly?) communicate extra context
"It was a very sunny day. [高兴]" <- same thing, but v1 pictos
"It was a very sunny day :)"
"It was a very sunny day :\"
"It was a very sunny day ;)"
Some web forums started replacing them with images which lead to designers inventing more emojis, and due to Japanese carriers wanting the same for SMS those got incorporated in character encodings.
What seems strange to me is that most emoji that exist are completely useless for the purpose of conveying emotion, or encoding any useful information that can't be expressed in a word. It's like some designer had to fulfill a quota or someone wanted to just "have more emojis". Yet the most popular emojis are clearly still used to convey emotion [1].
That is major advantage if that would prevent adoption as replacement for alphabet.
If they looked much different we wouldn't recognize them!