There are some real issues with English spelling, like the inconsistency of pronouncing 'ea' as /i/ or /ɛ/ (consider, uh, read and read). But 'ghoti' isn't one of them, because that's a case where there's not a lot of ambiguity in English pronunciation.
[1] The worst offenders in English pronunciation are when English borrows foreign words both with foreign pronunciations and foreign spellings.
Im ESL, I struggled with English spelling as much as the next latin speaker who's already learned to read and write in foreigner.
But now that I get the reason behind it, I love it. I consider English orthography worthy of UNESCO protection, even. In fact, I am annoyed at the regular spelling of my two latin languages that have left so much history behind.
Or natives. It is slower for children to learn to read English than other languages.
My kid took two and half years.
The Chinese take 10 years.
So what? Are the Chinese terribly educated?
Knowing etymology is a an easy way to memorize things.
You do have to have some familiarity with the source languages, but if it’s an unfamiliar but nativized word, those are almost always ultimately Latin or Greek.
You might also just happen to know a smattering (or even a lot) of Greek and Latin.
By all means a fun aspect about English that you can look at a word and guess the origin, and it's pretty satisfying to pull it up on google to see your high because you looked at spelling. This novelty has come at the expense of many other things that would have increased its utility.
I guess I'll add one thing to list benefits, this probably has resulted in different dialects of English writing things the same way despite saying them differently. Singaporean English is very different from Scottish English, but the written form of the same statements for the most part decipherable by the other dialect.
"engage", "engorge", "engrave", "engross", "engulf" are all fairly common words that are either often or exclusively pronounced that way (some dictionaries might show /in-g/, but /n/ is really /ŋ/ before g or k, even if they remain). Since these can take prefixes, this also proves we're not limited to being at the start of a word. Searching for words that can be spelled with with "ing" or "eng" finds a few more but nothing super interesting (though a few are in the middle of a word).
Obviously words where "g" is pronounced /dʒ/ (like "j" for those who can't read IPA) aren't subject to this.
- engage - /ɪnˈɡeɪd͡ʒ/, /ɛnˈɡeɪd͡ʒ/
- engorge - /ɪnˈɡɔːdʒ/
- engross - /ɪnˈɡɹəʊs/, /ɪŋˈɡɹəʊs/, /ɛnˈɡɹoʊs/, /ɛŋˈɡɹoʊs/
- engulf - /ɪŋˈɡʌlf/
According to Wiktionary only engulf and engross also use /ŋ/.
English might make more sense if someone actually sat down and wrote out the real stress rules, rather than trying to cram everything into just "unstressed" and "stressed" and only caring within a word.
=====
"To" might be one of the syllables with the most possible stress levels, with at least 4 and possible more. As I spell them,
1. "too" - full stress. Common for "two" and "too", but possible for "to" under rare circumstances.
2. "to" - less emphasized but still arguably stressed; still has the "proper" vowel. Usually this is as strong as "to" gets; "two" and "too" often fall down to this level if before a stressed syllable. Arguably this could be split into "stressed but near words with even more stress" and "unstressed but still enunciated" (which occurs even within a register).
3. "tah/tuh" - unstressed, the vowel mutates toward the schwa. Very common for "to", but forbidden in a few contexts. May be slightly merged into the previous syllable. Can we split this?
4. "t'" - very unstressed vowel has basically disappeared; may or may not remain a separate syllable from the one that follows (should that be split?).
The infinitive particle can't be 3 (normally 2, not sure if 1) if the following verb is implied (but not if the speech is cut off). At the start of the sentence it also can't be 3, and 1 is possible as seen below though 2 remains the default. Note that many common verbs act specially when before an infinitive particle; although sometimes treated as phrasal verbs it would be silly to treat them as taking a bare infinitive as their argument.
Adverbial particle "to" when the phrasal verb takes a direct object can be 2 or 3; this likely depends on the specific verb it's part of. Note that many people parse this as a preposition (taking a prepositional object), but this is technically incorrect (though there are some verbs where it really is unclear even when doing the rearrangement and translation/synonym tests).
Adverbial particle "to" when the phrasal verb does not have a direct object is usually 2 or even 1 (e.g. in the imperative). Some heretics have started calling this a preposition too (unfortunately, often in ESL contexts), but this should be avoided at all costs; they're just too cowardly to give particles the respect they deserve. Probably the only common example in modern English is "come to", but there are several others in jargon or archaic English.
Particle/preposition (the parsing is arguable) "to" used between numbers (range, ratio, exponentiation, time before the hour) tends to be 3, especially if one of the numbers is a "two". With variables it is slightly more likely to be 2.
Preposition "to" meaning "direction", or "contact", or "comparison/containment" tends to be 2, but can usually fall to 3 (less likely at the start of a sentence, and can also be prevented by what precedes it, e.g. "look to" can fall to 3 without much effort, but "looked to" strongly stays at 2). Contrast with "toward" of related meaning, which takes effort to get from 4 to 3.
Preposition "to" meaning "according to", "degree", or "target" (including but not limited to the explicit expression of an indirect object with most verbs, which we could argue should count as a particle instead. If you're wondering what verbs are excepted, one is "ask" - it can only use "of", as in "ask a question of him") is much more strongly 2, and requires significant effort to force it down to 3.
Adverb "to" is always 2 I think, but this is rare enough that I'm not sure.
=====
"To be or not to be", as famous as it is, has a pretty unusual stress pattern for most of its words: full stress on the first "to", semi-stress on the first "be", no stress (but still full length) on "or" (normal), full stress on "not", some stress on the second "to", and some stress on the second "be" (more than "to" but less than "not").
Spanish is totally systematic in this sense and once you can read it, you can pronounce it.
English is a bit messy regarding to this, for whatever reasons.
Even among major languages, English isn't anywhere near the worst offender of copulating with other languages for features--it never really adopted foreign grammar, the way you see with, e.g., Turkic languages.
The creolization is why English has a relatively simple grammar, and all the word sources is why we have like 16-20 vowel sounds trying to cram into latin characters.
You mean "relatively simple morphology". English phonology and syntax are not simple at all (e.g. lots of information carried by word order).
I’m sure you have a solid basis for saying this but it’s basically impossible to write many sentences without by accident using French down to the original spelling.
I was going to highlight all the examples I used by accident myself in this post but I gave up because the links were making it too long.
This is why something like Anglish even exists https://en.wikipedia.org/wiki/Linguistic_purism_in_English
Note that the prevalence of native words in German is the result of a modern reform movement, not something that happened naturally within the language.
> [English] never really adopted foreign grammar
There's the argument that do-support is borrowed from Celtic.
~26% Germanic
~29% Latin
~29% French
~16% Other
RobWords covers this really well: https://youtu.be/PCE4C9GvqI0?si=4Wd6NFus4v1YqmC3
is there no accent variation in Spanish?
Such a 1:1 system would never work in English, because the way words are pronounced can be very different in e.g. Melbourne, Newcastle-upon-Tyne and Boston, for example.
In contrast, English has a deep orthography, where historical layers (e.g. Norman French, Old Norse, Latin borrowings) and sound changes (like the Great Vowel Shift) have led to a chaotic mapping between spelling and pronunciation. A consistent system wouldn't eliminate dialectal variation, but it could reduce ambiguity and aid literacy, as evidenced by languages like Finnish or Korean.
In Argentina: "autito"
In Colombia: "autico"
In Spain: "autillo"
the same rule applies for all words, not only for cars.
-ito it's almost the universal way everywhere in the Hispanic world.
-ico it's widely used in the South of Navarre and Aragón and everyone will understand you. Heck, it's the diminutive from used by the hick people, and thus, it's uber known, altough you might look like a bumfuck village redneck sheepherd with a beret by using -ico outside of Navarre/Aragón.
-illo it's more from the South, but, again, understood everywhere.
"ico" is used in many countries of Central America and Caribe. I asked someone from Colombia, so I'm sure about Colombia but I'm no sure about every other country.
Is "illo" used in Madrid? I think I heard it in movies or TV programs from Spain.
Valencian has 7 sounds though, two for e and two for o. Similarly, Catalan also (and in some circumstances the o sounds as u, when the stress is not in it and other stuff). But they still have quite strict rules.
Now, you can (and should!) accuse me of cherry-picking examples, since the rules are less consistent and/or vastly more complicated than what I represented. But I maintain that there are orders of magnitude more ways to represent vowel sounds than 5, and the clue is the context. Not, as many will suggest, memorizing each individual case (though there's certainly plenty of that going around, much like Spanish's infamous irregularly verb conjugations), but understanding categories and families and patterns.
English sounds usually are best understood with groups of three letters, rather than one letter at a time. If you looks at throuples, you'll likely find far more of that consistency we all so deeply desire.
The first two are not productive now in normal Spanish words: they are only used in old spellings that have irregularly been retained, and in loanwords from indigenous languages. But they do exist.
Xenofobia is an s, yes, and excursión is "ks" In fsct, Méjico is the traditional way to write Mexico in Spanish grom Spain until it was accepted the other form a few years ago. I still write "Méjico" myself.
And anyway, as you point out, even in Spain the form México is accepted now.
After all, it is where they come from originally and have their own spelling (colour vs color, etc.)
An x in standard spanish has always been the two sounds I told you and that mexican deviation is specific to Mexico.
Yes, it is over 100 million speakers but I was still assuming the root language in its original place as the reference. Sorry if I did not express it correctly.
The "root language spoken in its original place" absolutely did pronounce X like modern J.
"ll" in standard spanish is a strong english "y".
However, in spanish argentinian from the area of Buenos Aires (but not the argentinian Córdoba, which sounds more like colombian spanish) it is "sh", being that s something like a mix in-between of "j" and "s" + h as in "she" but the sound is a bit different.
Without being able to record some sound I cannot express it better but I am sure you can find something around. Javier Milei, the president, has such an accent.
In the last 40 years I've spent mostly in the USA I rarely have heard Uruguayan/Argentinian Spanish in person or in media, but was surprised to hear Messi and others in recent interviews use SH as in "puSH" for the Y/LL, this apparent has been a generational shift in that area, first in Argentina and then Uruguay. I'd sound old-fashioned if I were to go back to Montevideo these days.
In any case, her point wasn't to give a lecture on linguistics, but to impress upon the parents how complicated English really is to learn to read.
- /k/ can be written both c and qu, and k where it occasionally appears in the language (e.g. kilo) - and the u in qu is silent.
- /s/ can be written c, s, and z, though stress rules are different for c and z.
- r and rr are distinct sounds but r = rr at the beginning of words, I think.
- At least in Mexican Spanish: The "ua" sound can be spelled ua or oa (e.g. Michoacan, Oaxaca) - and also the breathy sound of j can also be written with an x.
- d has a sound a little like English voiced-th at the end of words (e.g. juventud)
The stress rules, to the best of my knowledge, is very systemaic (not 100% but I would say "almost" at least for the words in use). Even the stress rules are very uniform.
> r and rr are distinct sounds but r = rr at the beginning of words, I think.
This is still systematic reading. At the start of a word it is the strong one, yes. And when it is preceded by a consonant, such as in "enredar" (that is strong r). There is no exception of any kind here.
> d has a sound a little like English voiced-th at the end of words (e.g. juventud)
That is some dialects in some areas. We pronounce a clean d at the end in my area (around Valencia). It is also the correct, standard way to do it for spanish. The other is a deviation existing in León, for example.
This isn't entirely correct. A distinct sound that the mouth makes is a "phone". A phoneme is almost always a group of several phones - allophones - that native language speakers perceive as a single sound. Another way to phrase it is that if you change one phoneme to another one, it makes a different word (possibly a non-existing one, but regardless the native speakers would consider it distinct), but changing from one phone to another doesn't change the word.
For example, in English, the phoneme /t/ has allophones [t], [tʰ], [ɾ], or [ʔ] depending on context. OTOH [ɾ] is a distinct phoneme in Spanish, and [ʔ] is a distinct phoneme in Arabic.
Unfortunately these two are often confused, so one should be careful with such counts and comparing them - it's not uncommon when people count phonemes in their native language, but phones in other languages (when those phones sound distinct to them).
This can also vary significantly from dialect to dialect, since one very common thing in language evolution is for two similar phonemes to collapse into a single one while retaining the original distinction as allophones. For English, in particular, the number of phonemes varies a lot between American and British English (with the latter having more distinctions).
You’ve never seen the word before, but when reading it for the first time, you’ll probably pronounce it correctly.
English is awful, but French takes the crown on this one—though more because it has the same pronunciation for many different words and written forms.
English, on the other hand, the alphabet doesn’t map well.
Mood and flood both have “oo”, yet each is pronounced differently. You need to know the word beforehand to know exactly how it’s pronounced.
Is this not really the case, and therefore is French also guilty of having the same vowels/consonants pronounced differently for completely arb reasons?
I do not want to be offensive, there are lots more , but it is an amazing sh*tshow the mapping.
My personal favorite in English is "colonel" being pronounced the same as "kernel". Which is insane even from an etymological perspective because the word is a derivative of "column" (as in, a colonel is someone who commands/leads a column of soldiers).
Hungarian, however, is pronounced the way it is written, as its orthographic type is phonemic, whereas French and English are of type deep orthography.
Serbian is of the perfectly phonemic type. "Write as you speak, read as it is written" is a common saying.
IMHO purely phonemic orthography makes orthography unnecessary complex, as there are language features like assimilation[1] that happens naturally in spoken form but does not make sense in written form.
In contrast, morphophonemic orthography keeps systematic and consistent mapping between spoken and written form for individual morphemes, but not necessary for words, as in written form morphemes are just concatenated (to make words), while in spoken form there may be complex interactions.
[1] https://en.wikipedia.org/wiki/Assimilation_(linguistics)
If the variant get's too popular the two versions become the official spelling, for example "septiembre" and "setiembre" (September) are correct. I hate the second one and I never use it, but it's popular somewhere. After many years, sometime the old spelling disappears and is marked as archaic.
Why is Zhou pronounced that way?!
In general, it's not transliteration into English characters, it's transliteration into the Latin alphabet. That means that transliteration tends to be shared across the various European languages that use the Latin alphabet. And given that the English were one of the last powers to actually engage in the naval trade war, they're less likely to be the basis of a major transliteration effort.
In the case of the q and x, I believe it comes from 500-year old Portuguese.
Not just European languages. Pinyin is useful for everyone that has to interact with Chinese words, whether their first language is English, French, Swahili, or even Mandarin.
A lot of people might not realize that the primary users of Pinyin are Chinese people. The way typing Chinese works is that you type the pronunciation in Pinyin and then a box pops up with choices of characters from which you select the correct one. It's also used in dictionaries to give the pronunciation of unfamiliar characters.
Pinyin uses s in a very common way, z in the way of Italian, and c more or less in the manner of various Slavic languages. They are a sequence of related sounds: s is the fricative, z is affricated, and c is both affricated and aspirated.
Sh, zh, and ch are a sequence of sounds related to s, z, and c. Sh is a fricative articulated farther back in the mouth, zh is its affricated form, and ch is both affricated and aspirated.
And as a bonus, sh and ch match English usage, which isn't likely to have been a primary concern.
It's also worth noting that for many Chinese speakers, there is no difference between s/sh, z/zh, or c/ch.
(x, j, and q are what you get if you use the middle of your tongue, instead of the tip, to pronounce sh/zh/ch. They occur before front vowels; sh/zh/ch only appear before back (or central) vowels.)
A friend of mine remarked to me once that when she was in school, her teacher informed the class that English speakers would not understand what the pinyin letter "q" was supposed to mean, which I immediately confirmed. She thought this was hilarious.
What use is "q" as a letter at all in English? It makes a "k" sound and always occurs with a "u" after it. Why not use it for the "tch" sound? (Which, btb, is different than the "ch" sound.)
"C" is about the same -- by itself it always sounds exactly like "k" or "s". Why not use it for the "ts" sound?
As for "zhou" -- in English, z is very similar to an s, but voiced. So in pinyin, zh is just like ch, but voiced.
Lots of languages do this BTW. When people from Wycliffe want to translate a Bible into an obscure language without a writing system, they first have to invent a writing system. They could invent all new characters, but why? All it would do is make that language hard to type. So they take the sounds that language has, and map them onto Latin characters. Sometimes there's an obvious mapping, sometimes not.
Look up Welsh's spelling for another example of this.
What are you thinking of? There is no difference between those things.
But your major point here is correct; on the fundamentals there is no reason for the English alphabet to feature a Q.
> "C" is about the same -- by itself it always sounds exactly like "k" or "s". Why not use it for the "ts" sound?
With the modern alphabet there's no reason for a C either. However, the answer to "why not use it for the 'ts' sound" is pretty obvious - that sound isn't part of the English phonemic inventory. It occurs, but that is almost always just a result of what is supposed to be a bare /t/ being followed by /s/ for grammatical reasons. (For an example of the general feeling here, note that an English word cannot start with /ts/ at all.) Why would we use any letter to represent the "ts" sound? We represent it the same way it exists in our language, as a sequence of two unrelated sounds.
> So in pinyin, zh is just like ch, but voiced.
Technically the only voiced consonants in pinyin are m / n / ng / l / r. I think a voicing contrast was present in Middle Chinese, and there's one today in Shanghainese and presumably other Wu dialects, but not in Mandarin.
I'm talking about pinyin here. In Mandarin, there are to distinct sounds, one represented in pinyin by 'q', and one by 'ch'. It took me months to hear the difference, and months more to be able to pronounce them properly. I think there are other romanizations where the 'q' sound is represented "tch".
(In fact, I'm inclined to think that there are actually two different sounds in English as well; "witch" and "Charlie" don't feel the same in my mouth.)
> Technically the only voiced consonants in pinyin are m / n / ng / l / r.
I think we're using different definitions of "voiced". Other voiced / unvoiced pairs in English include g/k, b/p, v/f, z/s. See [1] for an "official" example of "voiced" being used the way I'm using it.
How else would you describe the difference between "qu" and "ju", or "chou" and "zhou"? The only difference I can feel is when your vocal cords turn on.
There aren't.
> I think there are other romanizations where the 'q' sound is represented "tch".
Well, maybe; there are a large number of romanizations of Mandarin. But there are no significant romanizations where that is true. It's q in pinyin, ch' in Wade-Giles, and ts' or k' in postal romanization.
> How else would you describe the difference between "qu" and "ju", or "chou" and "zhou"? The only difference I can feel is when your vocal cords turn on.
You could read my other comment in the thread. qu and chou are aspirated; ju and zhou aren't. Your vocal cords don't turn on at different points for those syllables. Mandarin Chinese doesn't use voicing contrasts.
> I think we're using different definitions of "voiced". Other voiced / unvoiced pairs in English include g/k, b/p, v/f, z/s. See [1] for an "official" example of "voiced" being used the way I'm using it.
Yes, I know what voicing is. You don't seem to know what consonants are used in Mandarin.
Compare https://en.wikipedia.org/wiki/Standard_Chinese_phonology#Con... .
So the idea here is that chou and zhou are related in a similar way that the t's in "top" and "stop" are related: your mouth and vocal cords are doing the same thing, but in one case you have the puff of air and the other you don't.
At any rate, going back to the original question: the logic behind the choice is still consistent. On this classification, in Mandarin, p and t and ch are aspirated, and in English p and t and ch are voiceless; b and d and j and zh are unaspirated, and in English b and d and j and z are voiced. (And q is mainly thrown in to fill the gap, but its pronunciation in English is voiceless as well.)
Or, to explicitly quote from the ref you shared:
> Such pairs [of aspirated and unaspirated plosives and fricatives] are represented in the pinyin system mostly using letters which in Romance languages generally denote voiceless/voiced pairs (for example [p] and [b]).
(But then you get Hindi with a four-way distinction, both voiced/unvoiced and aspirated/unaspirated in all possible combinations.)
They're spelled that way; I don't think they're supposed to be pronounced that way.
https://en.wikipedia.org/wiki/Aspirated_consonant#Voiced_con...
>> True aspirated voiced consonants, as opposed to murmured (breathy-voice) consonants such as the [bʱ], [dʱ], [ɡʱ] that are common among the languages of India, are extremely rare.
> Languages usually have either the voiced/unvoiced distinction as phonemic, or the aspirated/unaspirated distinction.
My understanding is that all of these options are fairly common:
- two-way contrast between aspirated and unaspirated
- two-way contrast between voiced and voiceless
- three-way contrast between voiceless aspirated, voiceless, and voiced
- three-way contrast for labial and alveolar stops; two-way contrast for velar stops
True, but most languages don't distinguish between [h] and [ɦ] to begin with, with one often the allophone of the other. So listening to Hindi it sounds like the same thing, more or less.
Yes, that makes sense -- I certainly learned something from this conversation. It makes sense that speakers would naturally tend to classify things along different lines, and in Chinese the aspirated / unaspirated classification makes sense.
That said, after having had some time to sit with the proposition that 'j' in the English name "Joe" is voiced, and the "zh" in Chinese word "zhou" is unvoiced, it continues to seem obviously false to me. It seems very much to me like mistaking of the map for the territory [1].
[1] https://en.wikipedia.org/wiki/Map%E2%80%93territory_relation
To determine the true nature of the phoneme in a given language, you need to "flip the bit" on voicing (importantly: without adding/removing aspiration!) and see whether native speakers will treat it as different or not.
> Hanyu Pinyin was designed by a group of mostly Chinese linguists, including Wang Li, Lu Zhiwei, Li Jinxi, Luo Changpei, as well as Zhou Youguang (1906–2017), an economist by trade, as part of a Chinese government project in the 1950s.
By the way, they are not “English” characters; they are Latin/Roman characters, and used in a huge number of languages with different spelling conventions. Pinyin was created for the entire world to use, not specifically English speakers.
"zh" is actually one of the more reasonable pinyin digraphs because it follows the same pattern as "sh". If "s" + "h" results in [ʃ], then logically "z" + "h" should result in [ʒ].
"c" is used the way pinyin uses it in many languages (e.g. pretty much all Slavic ones that use the Latin alphabet, for starters).
"x" and "q" are more questionable, but there's precedent for either in languages using Latin-based alphabets - "x" can be [ʃ] in Spanish, for example, and "q" is [c͡ç] in Albanian.
Note that the sound [ʒ] is common in Mandarin, but its pinyin spelling is "r". "zh" isn't voiced and is affricated.
Saying it has pronunciartion rules it is an strech. You have conventions.
In languages like spanish if you read a word, is very hard to misspronounce it.
It can't be an affrication, because /ʃ/ is not an affricate. (Although /tj/ is affricated, as /tʃ/ [think "gotcha"] - when you say 'ti', you're referring to words that were pronounced with /s/ rather than /t/.)
Wouldn't /sj/ -> /ʃ/ usually just be called "palatalization"?
(The specific phenomenon in the context of English appears to be called "yod-coalescence".)
Not really. There's no way to guess how many english words are pronounced based on the written form, unless you've heard it before. And of course the pronunciation may vary wildly based on region/country as well.
The most telling evidence of this is the existence of Spelling Bee competitions in english language countries. The fact that hearing a word being spoken is challenging enough to figure out how it is written that it is a competitive sport, says it all.
There are many languages where the concept of a spelling bee competition makes no sense at all, because as soon as you hear the word being spoken, it is 100% deterministically obvious how it is written. English, not so much.
But, french is much worse!
Nah. Having learned both, French is easier in this regard. It is not as random, it has rules they work most of the time.
Spelling bee is the opposite direction, going from pronunciation to spelling; not a fair comparison.
Because pronunciation rules exist, they're just never explicitly taught and instead learned through exposure. For example, here's someone reconstructing as many of the rules as they can: https://www.zompist.com/spell.html
French is funny to me because the written language and the spoken language are in some ways quite different, with written french introducing considerable complexity. aller, allait, allais, allaient, alleé, etc. Since the spoken context for all the conjugations is almost always clear, I'm not sure why someone introduced the extra complexity.
It's far from as bad as English, but here's a Reddit thread with lots of French words which are not spelt as they are written. Not esoteric words either; along the lines of hier and monsieur
https://www.reddit.com/r/French/comments/1269a2x/is_there_a_...
Whoa, very much not! I have spent the last 20 years trying to learn how to pronounce french words (my partner is a native french speaker, so I keep trying). The only somewhat consistent pattern I have is that the last few letters of each word are often silent, but even that is not really always consistent.
I'm fluent in 4 languages but french is an impossibly tough nut to crack for me.
https://en.wikipedia.org/wiki/Vuk_Karad%C5%BEi%C4%87#Linguis...
Spanish is also very predictable. While there are a few exceptions (like 'c' can be 'c' or 's'), they are very easy rules to follow, so never any surprises.
English and French are in the batshit crazy category. It's pretty much all random, you just have to know from memorization.
English is hard to both read and write.
> The most telling evidence of this is the existence of Spelling Bee competitions in english language countries. The fact that hearing a word being spoken is challenging enough to figure out how it is written that it is a competitive sport, says it all.
That's two exact opposite things.
Languages for which you know how to pronounce a word just from its written form => you can have spelling bee competition there.
Languages for which you know how to write a word when you hear it pronounced => no spelling bee competition.
I'll take French as an example : if you see "o", "au", "eau" in a word you know how to pronounce it. There is one and only way. But if you hear "o" in a word then good luck knowing how to write it. So you got dictées (spelling bees) even if you can easily guess how a written word sounds like. The existence of spelling bee competition in the English world is not proof that the language written word pronunciation are a guess.
But the fact that such words exist, in such large quantities that memorizing them all is so challenging that this becomes a competitive sport, is why engligh is so impossible.
And, like, I get it. We don't have a fully regular one. But this is like the people that think we don't have a single word to describe some things, when they have to basically ignore adjectives and many many synonyms to get to that idea.
Even better when folks complain that we have different ways to refer to people from other nations. Ignoring that a large part of that is that we heavily deferred to how said people wanted to be referred to.
The distinction is there. English can be used phonetically. We prefer to preserve the heritage of various loan words instead.
Hearing Americans pronounce the French loanword 'niche' as 'nitch' instead of 'neesh' is cringe-inducing.
English pronunciation is just kind of a mess (especially in the US). It is one of the few languages where highly educated mature people are regularly unsure of how to pronounce a word in their own language or where there is no agreed upon 'non-dialect'/standard pronunciation.
https://en.wikipedia.org/wiki/Anglosphere#/media/File:Anglos...
Which is worse, being unable to correctly pronounce a word (but still being close enough to be understandable) or being completely unable to write a word?
One that still gets me personally is "hyperbole"--I know how it's pronounced but when I read it, I still say "hyper-bowl" in my head more often than not. I don't think I've ever made the mistake while reading out loud to someone yet, but it will likely happen some day and when it does I will feel very stupid.
Well, here you go: https://www.merriam-webster.com/dictionary/niche#did-you-kno...
> I still say "hyper-bowl" in my head more often than not.
Same. This is where diacritics would fix the problem: Hyperbolé. Although hyperbolee would also work, of course.
This is definitely a problem when it surfaces. For example the Stormlight Archive [1] series has two voice actors narrating the audiobook, and they don't even agree between them how to pronounce half the made up names.
Fantasy novels predate the widespread popularity of audiobooks. It used to be quite expensive to distribute a large enough volume of audio. The old "books on tape" cost a lot of money, were frequently abridged, and only existed for the most popular titles.
https://twitter.com/andylevy/status/1506748105735159818 (not there anymore; maybe the account holder ditched Twitter)
TIL: gimp is gimp and not gimp? I always pronounced this like gin.
Yeah, that's what the creator said, and that's actually how I pronounce it, too. Just pointing out that "gi-" words can have both hard and soft Gs.
> TIL: gimp is gimp and not gimp? I always pronounced this like gin.
You learn something new every day!
Whoever says that English is a phonetic language does not know what a phonetic language is.
The property that characterizes a phonetic language is that you can properly pronounce a written word that you know nothing about.
Just as it would be silly to claim that Japanese is not phonetic. Of course spoken Japanese is phonetic. They even have two fully regular alphabets that can both express the same phonemes, but are used for different reasons. As well, they have a completely logographic set that does not relate to phonemes, even though it is used for most writing.
> Of course spoken Japanese is phonetic
"Phonetic" is not a feature of spoken language, but of the relation between other language forms (usually, written, but you could make the same distinction for, say, sign languages) and spoken language.
> They even have two fully regular alphabets
I assume from "two fully regular" you are referring to hiragana and katakana, but those are syllabaries, not alphabets. (Romaji is an alphabetic system, though, but I don't know where you'd find a second one.)
Fair that I should have said they have two phonetic writing systems, decidedly not alphabets. I'm not sure the distinction is one that matters for what we are covering here?
It's a feature linked to spoken languages, since it is a feature of the relation of non-spoken (usually written) language to a spoken language.
But it is not a feature of a spoken language.
> Sign language, for example, is not phonetic, as many users of it cannot speak or hear.
Yes, in causal terms, the fact many users of sign languages aren't familiar with the sounds of the spoken language is a reason sign languages tend not be phonetic, but they are not phonetic in definitional terms because the symbols in the sign language do not represent the sounds of spoken language.
But it would make no sense to call a spoken language phonetic (except maybe if it was a code for a different spoken language, in which the phonemes in one mapped to the individual phonemes, rather than ideas, of the other.)
I get what you are aiming at, but phonetics is about speech. Is why you can reliably say how many phonemes different languages have. If you had to cover all vocalizations that people could do, you would have a bit more trouble.
I'm largely comfortable with the idea that there is something lacking in the orthography of English. Fully comfortable, even. I'm growing frustrated with how many are pushing the idea that it is not phonetic. The system is literally to convey, in writing, the words that you would speak in English. And the word "phonetic" captures that perfectly.
If you want to argue that we are building a new use of the word "phonetic" applied to writing that supersedes "orthography" and related terms. You do you. It still seems nonsensical to me and only works if you ignore that we have an alphabet that is literally used to convey speech sounds.
There is "not being regular" and there is "not even trying, and getting it right by a stroke of luck from time to time".
I wish English was more phonetic. Spelling and pronunciations is a mess. However the language is mostly phonetic.
We Italians, when we were children, we were taught to read based on the written letters, and we were able to read any word. It was normal, during primary school, to pronounce a word correctly and then ask the teacher what it meant. This is something you can not do in English.
And the converse was true as well! An Italian child is able to hear the surname of a new acquaintance, or the name of the village they are from, and write it down properly. In Italian, the question "How do you spell it?" does not make any sense! Again, this is something you can not do in English. Nor can you do it in French, because in the past centuries they had ink to spare and as such they started writing down useless letters that they do not pronounce.
English is complicated because it's decentralized and there is no authority to regularize it. Which is a feature, not a bug.
1 - Being fluent in the national language does not prevent people from maintaining their dialects in parallel.
2 - Whether a language is phonetic has no relation to political issues concerning dialects.
3 - Whether a language is phonetic has no relation to whether people like to use it.
4 - English got decentralized starting with the Age of Sail, but the lack of correspondence between written and oral forms is systemic and older than that.
That's not really true -- there is and was a great deal of dialect diversity within England itself. It was widespread printing that allowed languages to be standardized at the scale of nation-states in the first place: the divergences that developed after the age of sail were reversing convergence that had only begun a couple of hundred years earlier.
And although versions of English from the south and east of England became the basis for modern standard English, other dialects persisted and sometimes spread around the world, so some of the differences between English dialects globally are due to disparate influences from different dialects originating within the British Isles.
e.g. Northern Italian languages are technically more closely related to Gallo-Romance languages from the other side of the Alps than to standard Italian.
In my experience learning Spanish, their loan words are Spanish-ized, being made to be pronounced and spelled in a format that makes more sense in Spanish. Whereas in English, the pronunciation and spelling is usually taken more directly from the source, so you get a bunch of instances where a word's spelling doesn't really match its pronunciation.
We're still taught very basic phonetic rules in English. Like how vowels have a long sound and a short sound, where "ee" is the long e sound, or "<vowel> <consonant> e" triggers the long sound for that vowel. But you're also taught that many words are exceptions (e.g. bear vs beard). And you learn there are patterns to the exceptions, like how "ea," if it doesn't sound like "ee," will sound like a short e, like in "head" or "breadth," and particularly in cases like "dream - dreamt" or "leap - leapt."
And if you do a lot of reading as a kid, you vaguely recognize in the back of your mind some words that seem to follow a different set of pronunciation rules not taught in school (e.g. rouge, mirage, entourage, entrée, matinée, parfait, buffet, memoir, soirée, patois), which you learn implicitly. I remember this as a kid, only later learning those were French.
And this lets you guess pretty well how you'd pronounce a word. Just with basic rules and a lot of input to learn from, you can guess how to pronounce pretty much anything with good accuracy, because there are rules, and even a logic to the exceptions, but the rules are overlapping, so it's more like a set of rules you choose from.
I'd liken it to machine learning. You can learn the rules without even being taught the rules, like I did in the case of French loan words. And there are probably rules we follow without even realizing it, just instinctively thinking it's the natural way to pronounce the word without knowing why.
I'm not saying it's as good as being as phonetic as Italian, but it's not like we just have to memorize the pronunciation and spelling of every word as though it were a structureless string of letters and a corresponding, unrelated sound.
Sorry for the long comment.
I guess I'm really confused. It's not like English is some Arabic language where the orthography is in a second nearly unintelligible languages? Or, Chinese or Egyptian hieroglyphs... ?
I'm arguing exactly what I wrote: a phonetic language is one when you can see a written word and pronounce it correctly, without knowing what it means and without having ever heard it before.
Edit - as an example, consider "door" and "pool": the written form is not sufficient to guess the sound to associate to the double o.
Being able to guess how something is pronounced sometimes is not enough to say that English is phonetically spelled. English often borrows spellings directly from the languages that it is borrowing a word from, those spellings are usually phonetic (based on the source language's rules), and due to the presence of certain peculiar sounds, one can often guess which phonetically-spelled language a word was borrowed from. That's not an English word being spelled phonetically, that's people being forced to become language detectives. You can get lucky and guess the pronunciation of a Chinese character that you've never seen before (based on the radicals), but no one would say that Chinese characters are a phonetic alphabet.
Other than the soundalikes "b" = "v" and in Latin America soft "c" and "z" = "s", when Spanish speakers don't know how to spell a word, it's because they are also saying the word wrong when they speak.
/s?
https://people.cs.georgetown.edu/nschneid/cosc272/f17/a1/cha...
A more direct phonetic writing system, like many other languages have, would make it much easier to learn how to read and write English.
Isn't the "o" in "women" stressed?
English is fucked up. The only way to learn how to speak it properly is by memorization.
Other languages like Spanish or Korean keep a near-perfect one to one correspondence between written form and expected pronunciation.
But maybe compare '-ough' in: cough, tough, dough, through, plough.
But ‘borough’ in British English doesn’t rhyme with dough. The ‘-ough’ is a schwa.
this is not hyperbole. Sure other places are diverse, however because of the unique nature of the US and its size it just ends up attracting and subsequently absorbing.
same pronunciation of sh in ship is found in
- sugar
- sure
- machine
- Chicago
- mustache
- sheikh
- nation (!!!!)
Can you notice that some of those words do not have any "s" in them?
English doesn't make any sense.
I pointed out the ship example from the text, which was used to demonstrate how "this early French influence over English, which arose from the Norman Conquest, is the beginning of the reason why English is written without accent marks. ... This was the French habit that the Normans brought to England: the use of extra letters to spell sounds that the alphabet didn’t have special letters for. This is why English has combinations like sh, th, ee, oo, ou that each make only a single sound."
That's an extra letter being used to indicate a different sound than the base sound, similar to how diacritics are used to indicate a different sound than the base sound ("the cedilla has the function of ensuring that a c can be pronounced like an s, despite coming before an a, o, or, u").
That's cool 'n all, but I believe that only applies to French writing in English for English people.
Many languages have combinations of letters that have a single sound, it's no excuse for not having accents.
In German one can write strasse and straße or müller and mueller (different writing, same sound). They too don't have accents, but words written differently also sound different: schon = "already" and schön = "beautiful".
But German, on one hand retained diacritic marks, on the other it's also almost deterministic about pronunciation.
a it's always /a/
ä it's always /ɛ/ or /ə/ like e
sch it's always /ʃ/ as in schule
ch it's always /x/ after a, o, u and /ç/ after e, i
and so on
English doesn't use diacritics, IMO, because English doesn't make sense, it's a pastiche of lowest common denominators, so fck diacritics, they are too hard, let's write words as we like and pronounce them the way we feel they should sound, regardless of how they are written.
But it could use accents, for example rècord and recòrd, present and presènt, pérmit and permìt it's just they never thought it could be useful...
Shrug. Yes, languages have different paths in their evolution. Film at 11.
I still like what this linguistics PhD wrote about the specific history of one aspect of English language evolution.
> English doesn't make sense
That is of course an exaggeration. Just because the rules are complex and full of exceptions doesn't mean there's no sense. Even if you reject all of linguistics, Shannon in “Prediction and entropy of printed English”, demonstrated that English is compressible, which means there must be some patterns.
Now to drink some maté.
That's exactly what "makes no sense" means, actually.
> demonstrated that English is compressible
of course it is
> which means there must be some patterns
Of course there are. Patterns are (almost) everywhere - even PI is normal, but not random - but patterns in English make little or no sense for a language born and developed among, in the same era and having close contact with, a lot of other much more regular languages. The two facts are orthogonal.
Even Sumerian is more regular than modern English...
You don't need an "excuse" for not having accents. Digraphs and diacritical marks are simply two different ways to mark a letter as being pronounced as "somewhat similar but different". Whether one is better than the other is a matter of subjective perception, and it's very common for languages to not do it consistently. For example, Spanish has "ll" but also "ñ" (ironically the latter used to be "nn"!), and Czech has "č" but also "ch".
What's criminal about English is not the lack of diacritics, but rather the extremely convoluted and hard to predict rules for interpreting digraphs and trigraphs. If "ch" always meant the same thing, it would be just fine.
about accents, see my examples. they are used to disambiguate, which is a bonus in itself.
> If "ch" always meant the same thing, it would be just fine.
that is my take too: in German you have ss and ß, for historical reasons, but both sound the same and have a predictable pronunciation, always.
Infinite/finite regularly related, too - the reason the pronunciation of the finite cluster changes is due to stress differences (initial in- always takes the stress, and then the following syllable must be destressed). Note that the long vowel at the end comes back in the 4 syllable "infinitum", again due to regular stress rules.