This must be the world's largest tech debt. :)))
This must be the world's largest tech debt. :)))
Obviously, any artificial language is going to be much simpler. But these have never caught on for a variety of reasons.
Spanish is very willing to accept new words, and as diverse as English in terms of decentralization. Grammar doesn't make or break a lingua franca, number of speakers does, which is where English really shines.
So what does the RAE do you ask? They write grammars and compile dictionaries, describe phonology and answer people's questions on Twitter. Just like what Merriam Webster or Oxford would do, but the RAE has official backing and creates consensus among the hispanophone countries. English is a regulated language, just not officially regulated.
Some other language authorities do not take a usage-evidence-based approach to defining their dictionaries, and take into account cultural or historical concerns.
The role of regulation is mostly a legal role employed by countries that use a civil legal system. Most English speaking countries use a common law legal system and so interpretation of words is fairly fluid and subject to interpretation by the courts. Civil law does not have this kind of flexibility and room for interpretation, so courts will make use of regulatory bodies to uniformly interpret language.
This is why almost all countries with a civil law system have a regulatory body, while countries with a common law system do not.
This doesn't mean that learning the language is easy--no language really is. And English has some things that make it harder to master, especially its very large vocabulary.
Some of it is covered in this recent episode: https://slate.com/podcasts/lexicon-valley/2021/03/english-la...
Adults are really good at learning new words, but struggle with morphological and gender systems. Complex morphology seems to have some benefits that push languages towards including them, but are only sustainable when those languages are predominantly learned by children, who can pick up morphological complexity easily, not adults, who can't.
You have no idea :)
Having spoken English for almost 30 years by now, I am still not sure if "X has died" or "X died" is correct, or in which context.
But the other direction would be worse. Languages like Czech, with its 20+ classes of declension, must be a true nightmare for any English native speaker to learn.
Okay, strictly speaking, this is a distinction in aspect, not tense. But colloquially, tense, aspect, and mood are all referred to as "tense", especially since the conflation is present in most Indo-European conjugation patterns.
"X has died" is the present perfect. The perfect aspect is kind of like the past tense in that it is referring to something that has happened. Indeed, the past perfect ("X had died") is usually described as "the past of the past". But in keeping the tense in the present, the present perfect means that the past occurrence has relevance to the present. This can carry a few connotations. It can be a recent past, especially if you use "just" as an infix (c.f., "X has just died"). Or it can highlight the consequences of the event having occurred (e.g., "Our lord has died. What will become of us now?"). In any case, the speaker is drawing the listener's attention to the connection between past and present when they use the perfect aspect.
So which is correct? "X died" you would expect to find more in a biographical context or maybe a novel. "X has died" would be common in a news report, or someone informing you that a loved one died not too long ago. Which is more correct in a given scenario can usually be informed by the dominant tense in surrounding text; after all "X has died" is the present tense, despite conveying an action that happened in the past. If there's not enough text to dictate a tense, then it's often the case that either form will end up being acceptable--it just sets the tense that will be used.
"Have you [ever] been to New York?" "Did you go to New York last week?"
"Have you seen Star Wars?" "Did you see Star Wars this afternoon?"
Might not work in every circumstance but a good rule of thumb.
Vocabulary isn't a problem IMHO.
The fact that I have to learn each word twice (how to write it and how to say it) - is.
I was learning German for 3 years at school. After the first month I had no problems with pronunciation. Now after almost 2 decades of not using it I can still pronounce any German word I see.
I've been learning English since I was 10 or so. I'm 36 now. I still have many English words I know (and use correctly in writing) that I'm not sure how to pronounce.
Do you mind sharing some examples? As a native English speaker, I'm so curious! Do you think that if you heard them without seeing the word you'd realize what the written form was? Or might there be words where you know the written form and the spoken form and don't realize it's the same word?
There are quite a few others that tripped me up over the years like "cleanliness".
In general, "Chaos" aka "Dearest creature in creation" shows this problem (I would still struggle to read it even if I know every word there): https://pages.hep.wisc.edu/~jnb/charivarius.html
Parallel - for some reason the second a is eh not ah. Can't remember that, have to check it every time.
I play a lot of D&D over the internet in English and even as common word as "sword" is for some reason hard to remember. Every time I have to guess if the "w" is pronounced or not.
been == bin ? - the rules for that are just evil
> Do you think that if you heard them without seeing the word you'd realize what the written form was?
Sure, from the context if not instantly. I listen to a lot of English media with different accents (I watched the whole Big Bang and IT Crowd and I listen to Critical Role when I'm commuting).
> might there be words where you know the written form and the spoken form and don't realize it's the same word?
Leicester and queue. But these are famous enough that I remember them now. I obviously won't be able to give you examples that I still haven't realized ;)
> for some reason the second a is eh not ah
> been == bin ? - the rules for that are just evil
Actually, the rules are rather simple. They all have to do with unstressed syllables in English: unstressed vowels are reduced to /ə/ or /ɪ/ (the latter is what comes to play in your been -> bin use). Stress rules in English are not simple compared to other languages, and I can definitely see where non-native speakers might get confused.
One downside of the schwa reduction rules is that it can trip you up when you realize that you need to spell a word with a reduced vowel and you're not sure how it's actually written, because every vowel can be reduced to /ə/.
But the best examples are in the poem _The Chaos_ (http://ncf.idallen.com/english.html).
Through, though, throw, tough... I know a trough exists but I have no idea how it's pronounced.
Trough is /tɹɔf/ and rhymes with "cough".
See https://pl.wikipedia.org/wiki/Og%C3%B3lnopolskie_Dyktando
But the mapping sounds->letters isn't as obvious, because there's some tech debt there (ch == h, ó == u, rz == ż or sz depending on the preceding letter).
So if you know how the word sounds you're not always sure how it's spelled, but if you know how it's spelled you always know how it sounds.
In English the mapping is non-obvious both ways.
For example I'd seen adenovirus written down, but never heard it said out loud, I was describing the vaccine I'd had to friends, one of whom works in medicine (a doctor, but not of medicine) and she corrected my pronunciation because she's used that word plenty of times so she (presumably) knows how to say it correctly.
Even for a completely native immersed speaker, there's just no clue in English how to correctly say a completely new word you've only seen written down, so you're at no disadvantage there. For "real" words there may be an etymological clue, but those aren't reliable. In fiction it's anything goes. Hearing fictional words I've read pronounced out loud in movies is as weird for me as seeing the (inevitable) transformation of a woman described as plain in the books into a beautiful Holywood actress...
It's obviously a bigger problem with some common English words - either where they are actually two separate words with different pronunciations but the same spelling, or worse, one word but with different stress patterns. But once you've got a fair-sized vocab the new words you're learning won't have that sort of weirdness.
It's definitely true that if you're not confident pronunciation can really be an obstacle, fortunately the huge vocab helps again - a (non-English native but UK citizen) friend of mine will carefully choose to talk about liking the "seaside" never the "beach" because she's concerned she'll manage to make people think she said "bitch". She has a few other words like that, in each case English provides convenient alternatives.
Is her native language Spanish?
It's really cool how you can have "blind spots" depending on your native language. To me, the difference between beach and bitch is huge, because my native language uses short and long vowels extensively, and there are tons of words that only differ in a single vowel length.
But at the same time, I have other blind spots in English. For example, I have to make an effort to remember to use sounding "s" and "j" where appropriate, and the lack of those is a dead give-away for identifying Swedish English speakers.
I don't know why but a lot of people my age do worry about it and get very uncomfortable guessing at a pronunciation of a new word. Doubly so for names. I guess it's an insecurity thing? I have no problem just going for it, with a little question tagged on or just an upward tone if I'm really unsure.
Thanks for that anecdote, I somehow thought that I am alone living with fear of that happening :)
Written Finnish is mainly a written thing, and while there is a close correspondence between the orthography and how it would be read aloud, when Finns actually speak they use spoken Finnish (puhekieli), which isn't standardized and varies from region to region.
When reading Finnish, one pronounces the words very closely to how the word is spelled, regardless of whether when they speak they do so in an altogether different manner.
An example of a language that has it worse than English is Tibetan, which hasn't had a script reform since 800: https://www.youtube.com/watch?v=btn0-Vce5ug
I think the fact that it is a lingua franca is one of the main reasons keeping any spelling reform from occurring actually. It's in far too wide-spread use and there isn't any centralize authority that would do the spelling reform. Maybe take solace in the likely fact that it written English and spoken English will probably only be _more_ different as time goes forward. In other words, you have it easier than all future generations.
Also as much as spelling matching pronunciation is a convenience, it isn't really necessary. The variety of spoken Chinese languages using the same characters is greater than the spoken romance languages. Maybe English really slowly becoming more character-like over time. There are languages that can undergo spelling reforms, and there are languages people actually use.
There there is a very real cost to not mapping spoken and written languages: kids need to spend more time in school learning basic reading skill - time that could be spent either learning something else, or playing.
As I recall it speaks about the desire by the Japanese leadership of the time to build AI translation and related it to the challenge of full literacy in Japanese due to the requirement of learning 4 alphabets. Hiragana, Katakana, Kanji and Romanji.
Of course, it could have been worse. We could have ended up with French as the lingua franca (yes, I know what franca means) where there is almost no correlation between written and spoken language.
Going from spelling to pronunciation in French follows (admittedly complex) rules that are rarely broken except for common words (or endings such as -ent). Vowel pronunciations for a given spelling are far more variable in English, and often depend on the etymology of the word. Plus, English has word-level stress that is not marked in writing (French has none, and it's marked in Spanish), and moving the stress will usually make a word unintelligible! That alone makes writing => pronunciation very difficult.
Unsurprisingly, we can vaguely quantify this by looking at dyslexia amongst languages. English and various Southeast Asian languages that rely on Chinese ideographs are by far the worst, followed by things like Arabic, French, Hebrew, and German that have fewer exceptions but less guidance, and then followed last by things like Spanish, Cherokee, and so on that are truly one-to-one.
There are a number of ways currently used, but I have a new one to propose: compare the size of two G2P models (1 for each language), which have similar RMS errors. Assuming they are generated using similar techniques, the one which requires the bigger model probably has a less clean phoneme-to-grapheme correspondence.
Even ignoring all of these, its clearly not bijective. For example:
C --> /k/, /θ/
Z --> /θ/ [0]
K --> /k/
Q --> /k/
G --> /ɡ/, /x/
J --> /x/
N --> /n/ (with several distinct secondary articulations), /m/ (rarely)
M --> /m/
R --> Can be tapped or trilled.
Etc. You can go here and see many bijection-failures here: [1]
I am being intentionally unfair to Spanish (which truly does have a much, much better phoneme-grapheme correspondence than English[2]), mostly to illustrate the point that there aren't really any languages which have a 1:1 mapping between spellings and pronunciations. Even if you decide to use the IPA to write your language, non-standard dialects end up needing to read words that don't match their pronunciations. What happens when inevitably the language undergoes change - do we update all of the books to use the 'new' spellings of words?
The ideal orthography shouldn't be completely 1:1, but it should be relatively shallow. From that perspective, Spanish orthography is a fairly attractive option.
[0] The non-1:1 situation with /θ/ gets much worse in most dialects of Spanish, where it is not distinguished from /s/. See: https://en.wikipedia.org/wiki/Phonological_history_of_Spanis...
[1] https://en.wikipedia.org/wiki/Spanish_orthography#Alphabet_i...
[2] Look at how effective Spanish-speakers are at reading without "decoding" compared with Portuguese, which also has a good p-g correspondence. In particular, look how much faster the Spanish students are at pseudowords, on page 141: https://www.academia.edu/17872463/Differences_in_reading_acq...
The function from written Spanish to spoken Spanish (provided we are talking about a single dialect) is surjective, but darn close to bijective, especially if we exclude words of recent foreign origin.
> Many people expect … to predict the spelling from the pronunciations-- not realizing that few orthographies meet this goal. It's far from true of Spanish, for instance, which is often held up as an example of a good orthography. I stopped fervently admiring Spanish orthography when I saw a sign in a Mexican bakery with about one spelling mistake every third word.
So, no, hardly bijective!
I do, however, wish natives English speakers were more aware of this for their own sake. Most seem to default pronouncing vowels in foreign words as if it was English, whereas they'd be much closer to the correct pronunciation of they defaulted to pronouncing it like words for any language other than English they might know even a little of. To me this should be one of the first things you learn when learning your second language as a natives English speaker. It even holds true for romanizations of Asian languages like pinyin for Chinese or Hepburn for Japanese