Because of that, my grand father and his sister happened to have different family name in their ID cards: when they got registered by the priest at the beginning of 1900, on a paper book of course, the priest who registered the sister added an "I" at the end of the family name as he thought that that should be the right spelling in proper Italian. At that time ID cards did not even exist so nobody bothered.
e.g Papadopoulou vs Papadopoulos.
Then there are the cases where a Greek wife takes the male surname form to co-exist more easily in other countries where it is expected for a husband/wife surname to match (or if different, be more considerably different).
The only place where the difference between 1 letter and 2 letters still is visible is in names that start with a long ij, because of the capitalization. E.g. the name of the city IJmuiden is written like that and not as Ijmuiden.
In handwriting it's written as one character, but I think that's just a ligature: https://nl.wikipedia.org/wiki/IJ_(digraaf)#/media/Bestand:IJ...
Also, when spelling words people use 'ij' as a single letter for instance a Dutch person would spell 'mijn' as 'm', 'ij', 'n' (and they would likely not say 'lange ij' because there is no word 'mein' in Dutch.
It also causes no end of trouble with storing and searching because a Dutch person might expect IJ to be sorted after 'X' but instead it appears sorted as the combination I J . The fact that different reference works use different methods doesn't help either.
Not anyone under the age of fifty, unless they want to deal with the follow up question too. ("Is that a long or short ij/ei?")
Here's a screenshot from just one random Dutch children's spelling quiz online:
Separate letters for 'ij', despite it formally being a single letter.
Czech has this too, sort of, where "ch" is considered one letter but it is composed of two separate characters.
The language is unfamiliar to me, and quite intriguing. I find the pronunciation of Zuid very counter-intuitive.
Here is the full alphabet
A B C Č Ć D Dž Đ E F G H I J K L Lj M N Nj O P R S Š T U V Z Ž
More than you ever wanted to know about 'IJ':
https://en.wikipedia.org/wiki/IJ_(digraph)
I still have a typewriter that has it as a single letter.
Yes, this can lead to bugs in software, but so can anything related to names.
It is deprecated in Dutch to use a single codepoint for the IJ. That was never really an option in any of the character encodings in popular use.
The fact that it is one letter is relevant in cases like (vertical) lettering (which most designers nowadays fuck up), in typography (the number of fonts which make ij look awkward and unaligned is huge), and in collation and sorting using a Dutch locale. I will defend its proper use and treatment where possible, but representing it as a single codepoint is not a sensible goal, and never was.
I do not believe this to be correct. Wikipedia says it's a digraph of two letters. It does say that the codepoint is deprecated, but it's only defined as "compat", not as deprecated in the unicode data.
If you dig deeper on the Unicode website you'll find that the reason those codepoints are included is compatibility with 'certain very rare legacy (non-Unicode) character encodings'. They are not 'deprecated' as compatibility characters for those old legacy encodings, but 'deprecated' as suitable for rendering Dutch text unencumbered by those early code pages.
If you are claiming that 0x0132 is a codepoint in common use or required for correctly spelled Dutch, you are mistaken.
I made no comment in support or opposition of that. I cannot talk to that as I’m not familiar with either Dutch or that letter. However there are many compat characters in Unicode and there are incredibly few deprecated ones so I was addressing the deprecation claim (and what I believe is a misuse of the term letter). You’re interpreting things into my replies that just aren’t there.
Compat characters are very useful and even if they are not stored all the time, they often show up in text processing in memory for better glyph selection.
https://unicode.org/faq/casemap_charprop.html#6a
> The Unicode Standard encodes these two compatibility characters [0x0132 IJ and 0x0133 ij] to provide support for roundtrip conversion of the Dutch letter 'ij' in certain very rare legacy (non-Unicode) character encodings. It is strongly preferred (and far more common) to use the two character ASCII sequence 'ij' to represent this letter instead.
You can dig in the Unicode mailing lists for discussions on this from over twenty years ago. The bottom line is that you shouldn't use 0x0132 and 0x0133 in modern text. By now this is a resolved issue.
The 'IJ' is a letter, culturally speaking, in the sense that it is capitalized as one, and that it is often rendered as a single unit (which you can see in vertical lettering, if done right), and sorted as if it was a single letter. In terms of character encoding however, it is a 'i' followed by a 'j'.
Cf.:
* https://onzetaal.nl/taalloket/ij-plaats-in-alfabet (first sentence!)
* https://taaladvies.net/ij-alfabetisering/
See my other posts for 0x0132, which is a compatibility character for rare obsolete character encodings which did encode it on a single codepoint.