Orthographic Depth
en.wikipedia.org
en.wikipedia.org
(In the other direction, you may have to ask people to "write in the dialect", because it's so common to "clean up" text by removing the colloquial features of speech.)
The alternative, more common explanation is that there are two different registers of the language, one more suitable in writing and one more suitable in talking.
I wouldn't expect that to be the case at all, except for a small amount of intellectual, pedantic people. The vast majority of Finns reading kirjakieli aloud are going to subject it to many of the puhekieli sound changes. And of course, like with many other languages, if they speak a regional dialect with significant phonetic changes, they will apply those sounds instead of what is represented in writing.
> The alternative, more common explanation is that there are two different registers of the language…
Yes, exactly. The languages I gave as examples above, are typical examples of diglossia.
Then I'll be surprised. After all, that doesn't match the general understanding of Finnish having very shallow orthography (see e.g. the Wikipedia article). I'd be interested if you know research on this.
But even making all the sound changes does not turn the standard written language into the colloquial language, which entails some vocabulary and grammar changes as well.
Let me reformulate the experiment: Someone reads a text aloud as it's written and you ask if people still understood, if something was said incorrectly or even if they notised anything weird. It shouldn't be weird in any way, as you hear non-colloquial language all the time on TV etc.
> if they speak a regional dialect with significant phonetic changes, they will apply those sounds
Then I'll be very surprised. To begin with, wouldn't that in most cases go totally against the phoneme-grapheme correspondence of Finnish?
Secondly, when you intend something to be read aloud in a non-neutral way (colloquially or in a dialect), you write it that way to begin with. Why would the reader translate neutral language to non-neutral language and how would the reader know which dialect to use?
Finally, as the reader would be essentially translating the text from the standard language to the dialect, it would seem unnecessarily demanding for an otherwise simple task.
> The languages I gave as examples above, are typical examples of diglossia.
Finnish as a typical example of diglossia? The differences between the two registers are not that big to begin with, and it's a continuum with e.g. the languages spoken on TV falling mostly in some unspecific region in the middle.
I do understand that the prevalence of the two extremes can be problematic to learners of Finnish as a second language, because in a way you have to learn two dialects at the same time even though one is mainly written and the other is mainly spoken, and you never know which exact spot on the continuum is used by someone.
It can only be expected that someone reading a fossilized written standard aloud is going to apply various sound changes. When reading the written standard aloud with the standard pronunciation is seen as someone only newsreaders or solemn speech-givers do, then it will seem very stilted and pedantic to do so if you’re just an ordinary person. Now, the investigator could try to nudge the person into using standard pronunciation, but I would expect this to provoke some offense, because it might be suggesting that the person’s dialect or idiolect is wrong.
Moreover, there seems to be a widespread conception in Finland that actually sticking to the written standard in everyday writing, or pronouncing it aloud strictly according to the standard, is a marker of autism. Subconsciously, people try to show that they are more loosened up.
> Finnish as a typical example of diglossia? The differences between the two registers are not that big to begin with
Yes, I have often seen Finland described as an example of diglossia. The differences between the two registers are so great that foreign learners typically get 1.5 years of kirjakieli, and then they have to nearly start all over again with puhekieli using dedicated learning material. But it is not only foreigners: I have worked with Finns born abroad who learned only puhekieli from their parents, and upon moving to Finland and needing to be able to find jobs, they had to spend some time learning kirjakieli from a specialized programme, and by their own admission they found many features of it baffling at first.
> Let me reformulate the experiment
Your reformulated experiment only supports my claim of diglossia. Of course someone listening to a standard-language text read aloud will understand it and won’t find it weird. But it’s the same case for reading fusha aloud to a group of speakers of an Arabic dialect.
"Esperanto, Arabic, Finnish, Korean, Serbo-Croatian and Turkish are very shallow both to read and to write." Borne out in my anecdotal experience as a peace corps volunteer in very rural Korea, where universal literacy among very young children was taken for granted.
For example, "free of charge" is "besplatno". However, the first syllable is actually "bez-", with /z/ devoiced to /s/ by the last consonant of the cluster. The prefix "bez-" will appear undisturbed in other words, where the root starts with a vowel or a voiced consonant.
This is also done with loanwords, like "apsolutno" (absolute).
However, it comes from a general pronunciation pattern for consonant clusters, something that could be defined outside of the actual spelling. Why obscure the internal structure of a word like that?
It also exists in the opposite direction, by the way. Compare Serbo-Croatian "udžbenik" with the more elegant Slovenian "učbenik" (although Slovenian spelling has other, more serious problems than Serbo-Croatian, regarding some of the vowels).
"This means that the spelling reflects to some extent the underlying morphological structure of the words, not only their pronunciation. Hence different forms of a morpheme (minimum meaningful unit of language) are often spelt identically or similarly in spite of differences in their pronunciation."
https://en.wiktionary.org/wiki/immanis#Latin
https://en.wiktionary.org/wiki/alloquor#Latin
https://en.wiktionary.org/wiki/illumino#Latin
Most modern English words starting with a- or i- and a double consonant are inheriting or borrowing a pronunciation spelling from Latin that showed the Latin pronunciation of the ad- or in- prefix in context.
There's a challenge in transliterating Arabic because Arabic does what you suggest with the definite article al- <ال>. The /l/ assimilates in pronunciation to the following consonant if it is a so-called "sun letter".
https://en.wikipedia.org/wiki/Sun_and_moon_letters
So you could transliterate either al-shams (morphological or etymological or letter-by-letter) or ash-shams (pronunciation).
Why is that bad? Other Slavic languages also use the same approach. Consonant clusters usually sound a bit indistinct, so you have to make a choice which pronunciation you want to standardize for writing.
I'm a native Slavic speaker, and I know a bit of Serbian, and "besplatno" _does_ sound more closer to "s" to me.
> However, it comes from a general pronunciation pattern for consonant clusters, something that could be defined outside of the actual spelling. Why obscure the internal structure of a word like that?
To stay closer to phonetic accuracy. It's a trade-off. If you are a native speaker, you learn the spoken language first, so spelling being close to the actual phonetic pronunciation helps to learn writing.
Yes, it should be an /s/ for sure, but the spelling doesn't need to represent the /s/ sound in this case.
The /s/ sound could be spelled "z" here, and be represented by a rule that says "the voicedness of a consonant cluster is determined by its last consonant".
Serbo-Croatian phonetic spelling is like hardcoding stuff that shouldn't be hardcoded. And it focuses on the wrong stuff - why aren't there any accent markings in everyday writing?
Why? I don't get it. It's a very nice feature of Slavic writing: it's phonetic (at least in one direction). And it makes it MUCH easier for native speakers to learn compared to more 'hieroglyphic' languages.
> And it focuses on the wrong stuff - why aren't there any accent markings in everyday writing?
Because the stress pattern can usually be deduced, if you are a native speaker. So accent markers are really needed for exceptions or for non-native speakers.
And if we take another language's historic spelling as an example: Polish "rz" makes stuff easier for speakers of other Slavic languages, instead of spelling out the ż or sz sound verbatim.
... but there are constant dictation exercises (диктанты) in school to force children to memorize the thousands of exceptions in the orthographic rules of literary Russian. A written Russian sentence is (usually) easy to read - but the reverse task, writing down a spoken sentence, can be rather complicated for a child. Especially when the distinction between words and phonemes can seem rather arbitrary in a language full of clitics.
https://sergeytsvetkov-livejournal-com.translate.goog/105625...
> Especially when the distinction between words and phonemes can seem rather arbitrary in a language full of clitics.
Uhh... Russian doesn't have a lot of clitics. There are compound words like in German, but they are still understandable if you read them phonetically.
And most children learn to read first by sounding aloud individual letters, and then just training to do it faster. It's called "reading by syllables".
But Ukrainian is still even better. It's almost completely phonetic both ways and you usually can write down any word you hear without problems, even if you hear it for the first time. There are still some double consonants and a small amount of reduced vowels, but even they are still somewhat reflected phonetically.
https://youtu.be/sa3Tl3t88Mc https://www.youtube.com/shorts/82GgaT0Sx1o
I wonder how this would affect communities that speak AAVE for example.
[1] https://en.wikipedia.org/wiki/Swiss_German [2] https://en.wikipedia.org/wiki/Swiss_Standard_German
That has already happened, which is part of how we got into this mess.
See isle vs island.
IMHO The only real solution is to invent a time machine and keep England from getting conquered again and again and again. If you keep mashing languages together you are bound to get a mess.
https://en.wikipedia.org/wiki/Acad%C3%A9mie_Fran%C3%A7aise
This is how Turkey ended up with its shallow orthography: the government designed a new alphabet. Before that they used Arabic.
----
For example, in Year 1 that useless letter "c" would be dropped to be replased either by "k" or "s", and likewise "x" would no longer be part of the alphabet.
The only kase in which "c" would be retained would be the "ch" formation, which will be dealt with later.
Year 2 might reform "w" spelling, so that "which" and "one" would take the same konsonant, wile Year 3 might well abolish "y" replasing it with "i" and iear 4 might fiks the "g/j" anomali wonse and for all.
Jenerally, then, the improvement would kontinue iear bai iear with iear 5 doing awai with useless double konsonants, and iears 6-12 or so modifaiing vowlz and the rimeining voist and unvoist konsonants.
Bai iear 15 or sou, it wud fainali bi posibl tu meik ius ov thi ridandant letez "c", "y" and "x" -- bai now jast a memori in the maindz ov ould doderez -- tu riplais "ch", "sh", and "th" rispektivli.
Fainali, xen, aafte sam 20 iers ov orxogrefkl riform, wi wud hev a lojikl, kohirnt speling in ius xrewawt xe Ingliy-spiking werld.
I naturally read "iear" as "ear", not "year". And I read "awai" as something similar to "uh-why" (I'm having trouble translating the exact sound to text).
Other than that, I see no issues with the changes through year 5. It's not that hard to read and I'm sure I would quickly get used to it with practice. Past year 5, I don't understand what is happening and I literally can't read it.
It technically needs to be a different letter to mark a "short i", like "й" in Russian or Ukrainian. If we use "ï' for it, the "year" will become "ïear", "yellow" will become "ïellow". The "y" that reads as "i" will just switch to "i", and "my" will become "mai".
It's also completely impractical given the wide variety of very differently pronounced English dialects and accents. In other words, it requires massive pronunciation reform too. If that's on the table, why not alter everyone's pronunciation to fit the orthography...
And this is actually a loss. I have a fantastic edition of the Song of Roland with facing pages in Modern English and Medieval French, and there are nuances that would be lost in a straight translation to either Modern English or Modern French.
Updating grammar would be a giant pain in the neck. Saying "we're gonna write 'cat' as 'kat' now sounds a lot simpler. It would also greatly reduce the number of times I'd have to say things like "I'm about to use a word I've read a thousand times but never heard anyone pronounce, so forgive me if I'm butchering it." My kid was hugely into cats as a little kid, and it was cute when they wanted to tell us cool things about lee-oh-pards. It's less cute when I'm in a meeting talking about a whitepaper I've read about something I've never heard spoken.
At least when you encounter an unknown word in English you can more likely understand it even if you aren't sure how to pronounce it. Languages with periodic spelling adjustments to re-align spelling with (ever shifting) pronunciation strip this info away.
> Spelling in English is so hard we even have a contest for it, the spelling bee!
Pfui. In France there is (or was -- I no longer live there) a popular prime time spelling bee, not just for kids.
That probably seems like an intellectual concern divorced from day-to-day language use, but I think we take for granted how easy it is to figure out a word's meaning from its spelling once one has a grasp of written English. Words we typically think of as being clearly related, and pronounced as basically similar sounds are actually vastly different in basic pronunciation, and, thereby, reforming spelling in this way would eliminate visual relationships between related words that are otherwise totally non-obvious.
And all this to perform a phonological mapping that would be out of date almost as soon as it was settled on. To say nothing of the actual practical concern of arriving at a consensus on which phonemes in which words should be spelt which way, what of the absolute deluge of different accents to be dealt with for instance. That's before we get to the part where now English speakers lose access to all pre-existing literature after about 2 generations. The longer you look at the proposal, the more counter-productive it obviously is.
But reading? There are constant and consistent spelling reforms of Dutch, which partially aim to make it easier to read. My experience of it is that it's overwhelmingly regular (minus your odd melk surprise.) And apparently despite this, literature going back to 1994 empirically finds Dutch opaque for reading.
(Especially in the given context, where this specifically refers to "unvocalised Arabic", which implies that the written language doesn't include vowels.)
You can add diacritics to both Arabic and Hebrew to show all vowels, but this is rarely done outside religious texts.