How, btw, did this question follow from your point in the starting commentary?
I mean.. really?
I'm so astonished by that logic I don't even know what to tell you. Different languages can have different rules you know.
Edit: thanks for the downvote. How dare I tell a person who just claimed half of Europe is composing their alphabets wrong because Germans do it differently that they are short sighted.
Absolutely not. My argument was that Kazakh is far from unique in having ”extra letters”, which my native language’s alphabet also has.
On a personal level I far prefer latin-alphabet Slavic languages since it makes it easier to approximate an understandable pronunciation.
I of course agree with there not being a problem with extra letters because Polish does it without any significant problems as well.
No.
That said, letter with diacritic vs separate letter is a surprisingly fraught topic, the same character can be considered a letter in one language and not a letter in another. Look up collation rules in European languages for a real nightmare...
Is it Latin? Which version? See https://en.wikipedia.org/wiki/Latin_alphabet
Or is it all modern non-Cyrillic letters used in Europe that don't have diacritics?
In any case, diacritics are a positive. They're used to make a language as phonetic as possible, which is a net positive. As anyone can attest, learning to pronounce English words (or French, for that matter) is a lot of people wasting a lot of effort, which would be better spent elsewhere!
French (there are quite a few of them - think CA for example) has just as many idiosyncrasies as English when you really get in to it.
It is totally impossible to predict how an unfamiliar written word is pronounced in English, but quite easy in French (barring some ambiguous situations like if a noun ends in “s”.)
By the way, I don’t know what CA means. If you’re referring to the French word “ça”, its pronunciation is completely regular and predictable given the spelling.
[1]: https://en.wikipedia.org/wiki/French_orthography#Sound_to_sp...
[2]: https://en.wikipedia.org/wiki/English_orthography#Spelling_p...
It’s usually (not always) clear how a word is pronounced in French if you know its spelling.
The problem isn't Latin, the problem is how you use it. Nobody forced you to sometimes use C when you mean K. There are many languages where these letters don't overlap.
Nobody forced you to use both PH and F for the same sound. Why even use PH? It's not like anybody cares if that particular word was stolen from Greek or not. And if some people do - they can check it on wikipedia. Make it all use f and be done with it.
The whole too/two/to situation is ridiculous. Ridiculous spelling is also ridiculous. Why put "o" before only one of these "u"? Either put it before both or none... Also - beautiful. Really? Eau?
And why sometimes there's one "l" and sometimes there are two? Full, but plentiful. WTF?
If you bothered to refactor the language every century - English would be perfectly reasonable now. Instead the whole world is stuck with learning centuries of design debt to communicate.
I guess keeping compatibility over long periods of time can have a lot of value, for OSes and languages alike...
They've added 9 of their own to 33 of Russian Cyrillic for a total of 42 when using it for Kazakh. Now they're gonna add 6 diacritics and two digraphs, that's much less already.
Compare this to even Polish that has:
1. 9 additional letters: ą, ć, ę, ł, ń, ó, ś, ź, ż.
2. Technically no x, v nor q, but everyone knows their sound so casually or artistically you can use them, just not in 100% orthographically correct Polish.
3. 7 digraphs: ch, cz, dz, dź, dż, rz, sz.
And it's not really a problem in any way and the words are distinct enough that if you skip all the Polish specific stuff it's still 100% readable, people texting or writing online or Polish comments in code (in an ASCII file) often do that.
I've seen many systems (CJK, Arabic, Hebrew, various European ones) and have somewhat an interest in this stuff and I don't think it's a big deal at all.
I also think (as a layman) that it's really not about the system anymore but about fitting it to your language (within reason, e.g. don't try to write Polish with Chinese characters), someone could say Polish butchered Latin or gave it 16 warts but it works and I'd say the phonology is way saner than English (okay, that's not an achievement but still, no one tells English off for it's use of Latin despite it's crazy pronunciations).
I also wondered about an experiment of using both Cyrillic and Latin together at once in a language that already uses one of them to denote foreign words (e.g. Polish with foreign words in Cyrillic or Russian with foreign words in Latin), just like katakana in Japanese does (among other uses). I wonder if Kazakh writing (casual or maybe even official) with loanwords from Russian (if they exist, I don't know Kazakh, maybe they have none) might do that just for Russian foreign words, to avoid transcribing them into the Latin alphabet (which would be more awkward than just using them directly as they could in their Cyrillic which was a superset of Russian Cyrillic and would either ignore their own Latin phonology and diacritics or require a new transcription style that is unlike the other ones, thus adding to confusion).
Using Cyrillic and/or Latin of course has implications about politics, history and religion, with Cyrillic being associated with Russia, Soviet Bloc, Orthodox Christianity, etc.
However, the Turkic languages don't have a good start in any of these three, yet they're all written in one of them with a large number of diacritics.
From the linguistic point of view I posit, your primary concern in designing an orthography is: does this line up nicely with the phonology? If it does, congrats, your literacy rate just went up. Every digraph you add is another exception you have to explain, every sound with two letters is another exception you have to learn. My son is in kindergarten. He wrote "KUMIN" on a piece of paper and hung it on the door the other day. I asked him what it said, he told me "It says 'come in'". Was that obvious to you? It wasn't obvious to me.
So, in the grand scheme of things, I know a new Turkic orthography is probably a long shot. But wait, there is already a Turkic language with similar phonology: Turkish. Did they take the Turkish orthography and modify that? Not really. They had to come up with yet another g diacritic.
I hope they do normalize the spelling and make it phonetic. Then at least, they may improve literacy, assuming it wasn't as phonetic as it could be under the old alphabet. Because otherwise, what are they getting for what they're spending this huge amount of money replacing books, retraining teachers, fixing software, etc.?
In phonetic languages with digraphs there are rules to the exceptions. For example "sz" is related to "s" same way "cz" is related to "c". "I" makes the previous letter softer.
And there are Slavic languages with Latin script and almost no digraphs. Polish is kinda old school in that regard as it preserved sz and cz (funny thing - English name for Czech Republic uses Polish/Old Czech spelling).
English is one of the worst examples of using Latin script there is, and it could be drastically improved if it used regular digraphs instead of ad-hoc spelling. For comparison German from the same language family is much more regular and easier to pronounce.
The writing split in Slavic languages also follows closely some West/East (historical and contemporary) and Catholic/Orthodox (it depends, Czechs turned largely atheist by now from Catholicism but Poles, Russians, etc. mostly remain in their religions) splits more than anything, e.g. West Slavs have Latin and prevalence of Catholicism, East Slavs have Cyrillic and prevalence of Orthodoxy, South Slavs are a mix of the two and it shows, e.g. Croatia is more Catholic and more Western and it has the Latin alphabet, Serbia has Orthodoxy and both writings, Bulgaria has Orthodoxy and Cyrillic, etc.
Technically even Cyrillic which should fit Russian perfectly has things like И and Й which are two related sounds, because it really makes sense that a related sound has a similar letter with just a diacritic mark.
Even the Japanese hiragana and katakana have diacritics and digraphs, even though it's a fully bespoke (yes, based on Chinese characters and some Indian Buddhist scripts but they were modified and adapted very extensively and barely look like the originals from which they were evolved from and it was over a thousand years ago) system that evolved in Japan (an island nation), got a few official revisions to clear it up and is specific to only the Japanese language (and Okinawan and Ainu, but that's secondary and due to Japanese presence and didn't affect its design).
It's similar with Polish, for example ź is like z but with sort of wheezing the air under your tongue and through lower teeth instead of on top of it (sorry for the bad explanation, you can compare Polish letters and their bases on Google Translate or something and you'll hear the similarity of the sound).
There were attempts to Cyrillize Polish[0] and Latinize Belarusian[1] but they were shoddy at best and done by an occupants so that didn't go well.
The only funky thing in Polish that is better done in Cyrillic is digraphs (but then Russian has ть and we have a single letter ć, which is visible in base form of many verbs, like in Russian делать and Polish robić). Polish has cz and sz while Cyrillic actually has real letters for those sounds: Ч and Ш (Cyrillic also has Щ that sounds like sz and cz chained together like in szczęście but we don't consider that combination to be one letter or a quadgraph or anything like that). The digraphs also aren't a problem because they (IIRC, maybe there's some words and I can't think of it because I'm native) never appear like that naturally, if you get s before a z, it's a sz sound. Because of this it's not really an exception but a rule, if you gave a Pole a paper with just SZ or CZ or RZ on they'd do that one sound, not try to pronounce two letters.
Digraphs (in addition to j, w and y, leading to wafelek jagodowy by Ashens) might be one of the hardest things for learners actually because people try to hack they way through Polish using their language and they go with trying to slur the digraphs together while they are a different sound altogether from the two letters (it depends, like rz is ż, but sz or cz is like a soft or swishing s or c, but nothing of a z, sz and cz actually do sound like sh as in shoot and ch as in chain from English and they also have nothing to do with that h, then again h is sometimes silent like in honor and sometimes voiced like in holy..) that make it up so it sounds very off. They are also not that easy to pronounce, one Polish tongue twister is "w Szczebrzeszynie chrząszcz brzmi w trzcinie" and I once heard an African immigrant to Poland say how when he first came to Poland he felt like everyone is rustling and swishing all the time at him (which is actually accurate with regards to some diacritics), then again, his Polish was really good (he could pass for a native if he wanted IMO) so it's possible.
If not for digraphs then you could probably get away with learning pronunciation of each letter and then gluing/slurring those sounds together into a word. It's much easier than English where spelling is often only tangentially related to pronunciation and adding or taking away a letter somewhere else in the word can change pronunciation of other letters. There may be words that have something weird going on but I can't think of any right now (it might be my native bias though so be careful when hacking Polish that way).
Lots of people say that a spelling contest wouldn't work in their language and except for ó vs u, rz vs ż, ch vs h and some cases where it's hard to say if you heard ę or en, ą or om/on, c or dz at the end of a word (and maybe something else I'm now forgetting) etc. that's true in Polish too. If you say something in Polish a Pole can usually write it down on first try without even thinking about it. E.g. there is a few niche jokes about two fake useless devices called bulbulator and przyczłapnik, these two words are never used except for that joke but their spelling makes 100% sense and if someone heard the joke for the first time they could write it down too, no problem.
[0] - https://pl.wikipedia.org/wiki/Cyrylica#Cyrylica_polska
[1] - https://en.wikipedia.org/wiki/Belarusian_Latin_alphabet