"Any self-respecting" -- this is very, very hard. If you are Danish, your keyboard most likely has a way to enter Ø and you would expect Ø and O to be handled as two different letters. However, if an English article mentions the Øresund bridge, it's not unreasonable to expect a search on Oresund to match it. Or, in Hungarian, "kor" means period or age, "kór" means disease and "kör" means circle... but searching for "kor" in Chrome finds all of them. And Unicode won't save you: Ø does not have decomposition rules but ó and ö both do. So the common practice of decomposing - removing marks - recomposing route can either be too little or too much. The upcoming solution for Drupal will allow users to edit remove diacritics rules per language, no other way to do it.
This is Libreoffice taking that path, I know Firefox does the same: https://cgit.freedesktop.org/libreoffice/core/commit/?h=libr... see the icu::Transliterator::createInstance("NFD; [:M:] Remove; NFC", part.
In the last stage, a full DWIM-AI would do the matches. ;)
I once worked on what we called DYM tools (for "Did You Mean"). The goal was to assist native English speakers who were learning a second language find words they heard, in that language's electronic dictionary. We knew, for example, that native speakers of English have difficulty distinguishing the dental and retroflex consonants of Hindi, so the DYM allowed mismatches of those. A perfect spelling was at the top of the list returned from the dictionary, while a misspelling--such as writing a dental for a retroflex--resulted in the mismatched word being a little lower in the list. We tailored our DYMs for particular target languages (and always for native speakers of English). As you say, getting something like this to work for multiple languages at the same time would be difficult.
é is just a different accent of e, not a different character in the alphabet. ä is a completely different character in my alphabet. Similarly ö is a completely different character in my alphabet and should not be considered equal to o, if the text is Swedish.
For someone in the US or France, the ö or ë probably is just a pronounciation diacritic as in Raphaël or Cooöperation. Of course an American doing Ctrl+F on a website (Likely New Yorker!) would want to find "Coöperation" when searching for "Cooperation". Even weirder: when I go to the New Yorker and search for Cooperation I want to find "Coöperation" too!
So this is highly context dependent. Ideally I want collation/comparison to depend on the content of the page, not my browser/OS language.
Another idea would be to do that in Swedish text, "é" will always be represented in decomposed form, while "ä" will always be represented in precomposed form, so that you can tell the difference.
<blockquote lang="sv">...</blockquote>I'd even suggest browser to mess up formatting of non English letters (by using wildly different fonts) to encourage better semantic markup, but it is a bit hard-core and everyone would shout "compatibility breakage" at them :)
As long as it is user configurable which fonts to use for which language, I think that it does not break compatibility to do that. Actually, I think it is a bit good idea. The document should only specify the language and the style (e.g. bold, emphasis, normal, fix pitch, heading, etc) and then that combination is mapped to a font in the browser. (If the user has enabled use of CSS fonts, and such fonts are specified, then they would override those specified by the user. If the user has not enabled use of CSS fonts, then the user's fonts are always used.) This would be needed anyways due to the Han unification that Unicode does, anyways (and Unicode is very messy, anyways). (I mentioned before that Unicode can be good for searching, and if that is what you are doing with Unicode rather than for writing and displaying documents, then Han unification is probably desirable, although again the Duocode that I mentioned before may help even more.)
The correct way to handle this is by tagging the text with a language tag, as has been mentioned in other replies to the parent post.
Of course all of this is locale-dependent and it might be acceptable for ease of use to match "ä" to "a" for international users. But not in a german locale.
Proper I18N is very hard, and supporting Unicode is just a first step. Converting text between lower/upper/title-case can only be done knowing at least the locale of the text, and in rare cases a case-folding-roundtrip can even change the meaning of the text, so you will have to be able to speak the language.
This is unworkable, since "Busse" (several autobuses) should not match "Buße" (repentance). But "Busse" (incorrect spelling of "Buße" due to character set limitations/Swiss German spelling I believe) should match "Buße" by your argument.
Anyway, modern Firefox lets you choose whether to match "Apfel" against "Äpfel".
But: "Busse" is not a always incorrect spelling, it's just an "emergency" one, if you really can't use "ß" for some reason. One example is allcaps text, here "BUSSE" and "BUSSE" are not distinguishable, although they have different meanings. Also, there is Swiss German, where replacing ß with ss is the normal form and not at all "accepted in an emergency" like in German German: https://www.galaxus.de/de/page/schweizerhochdeutsch-fuer-anf...
Incidentially this is one of the instances where you really have to know the language and the context to be able to do a lowercase -> uppercase -> lowercase roundtrip, otherwise you might screw up with "Buße -> BUSSE -> Busse", changing the meaning. Or not, if the current locale is de_CH instead of de_DE.
Together, since very often search will be case-insensitive, I think that while you may be correct, being as strict as to not match here would not be what the user would expect.
Yes, again, I18N is hard and sometimes it is impossible to do correctly for a machine.
Szene should match ß?
Also, in this case, Sz is pronounced as the two distinct letters it is composed of. Whereas when "sz" replaces "ß", the pronounciation is just very similar to a plain "s".
Admittedly German isn't my first language. But it seems odd to be permissive in one case, but more restrictive in the other. In my experience there are two cases for searching: exact matches, like you'd expect an editor's find and replace, and more permissive searches, where you might expect to skip through several matches to find what you're looking for. In this second case, what's the advantage of being more exclusive in one case, then less exclusive in the other? Why not have a more permissive default (as mentioned by the other user's reply regarding Firefox's configurable search), even if it's technically incorrect?
Whereas for umlauts ä, ö and ü, the forms with ae, oe and ue are always "emergency replacements", even the Swiss use äöü very frequently (though not always). Also, there are almost never ambiguities when replacing ä, ö, ü -> ae, oe, ue. Yet there are very frequent ambiguities when replacing ä, ö, ü -> a, o, u, because singular/plural, conjunctive and other forms are derived that way and sometimes there are just different words that only differ in one vowel being an umlaut.
So there is a difference in the chance of being wrong when picking one form of a word for another possibly equivalent one. When you do ss == ß, you are often correct. When you do ä, ö, ü == ae, oe, ue, you are almost always correct. When you do ä, ö, ü == a, o, u, you are almost always wrong.
requires, not just "allows".
> even the Swiss use äöü very frequently
I can't think of any case where Germans would use one of these and German speaking Swiss would not. But as another Swiss peculiarity, some of our words contain "üe", e.g. "Üetliberg", "gmüetlich".
But, regex is a separate issue than languages of text.
...and now that I've complained about it I notice that there are extensions to enable regex search in the browser itself, nice
Also, as a Finnish user, I would not be impressed by a search for "talli" (stables) matching "tälli" (blow), or "länteen" (westwards) matching "lanteen" ('of the hip').
In the US, there's starting to be more push from people with ñ in their name to get support for the correct letter, but most government systems have zero support for anything but ASCII all capitals (sorry McWhoever).
On mine (77.0.1) it finds the OP's "The German character ä" as well as yours. Conversely, searching for "The German character ä" finds the OP's original and your "a" version. There is a "match diacritics" button for enabling/disabling this behavior.
Normalisation and - except for the soft-hyphen - ignoring of ignorable whitespace characters (ZWNJ, ZWJ, WJ etc.) on the other hand are still missing in Firefox (https://bugzilla.mozilla.org/show_bug.cgi?id=640856).
The number of times, in actual real life, that the inconvenience of the ambiguity outweighs the convenience of the overlap are not many.
YLMV.
Why? That is horrific! They are different letters. Why would you ever need that?
I have no idea. I don't speak Greek, and I am not familiar enough with its alphabet. The question, I suppose, is if there's a reasonable/intuitive enough mapping between the Greek and the Latin alphabet that it would be possible to search for Greek words using a non Greek keyboard. If yes, then I suppose my answer would also be yes - but I'm happy to be corrected by someone who is actually speaks Greek.
My point is that when it comes to things like searching, usefulness trumps purity. I search because I want to find shit. Not because I want to get perfect feedback on linguistic details from my software.
Usefulness of such behavior is objectionable. Characters with and without diacritics are different graphemes, part of different words. If i search something, i do not want to get plenty of irrelevant results. Just that happened to me few days ago, when i entered a rare word root to search box in Firefox and was surprised that got plenty of irrelevant results because there was common word differing in diacritic marks.
Even if sometimes there is a text written without diacritic marks where there should be, it is usually consistent for whole document, so if i search some word that contain diacritics, i know whether to enter the word with or without it.
I can imagine cases where enough context is available and you can only do the actually usable search. Heck, even Google doesn't give you any option, though I frequently hate that.
If you're going to have a "smart" search that decides to mesh these together, the option to turn off the "smartness" should be right there.
However, I cannot get it to find my friend Nikolay when I type the whole thing, as it doesn't match the final Cyrillic letter with my "y".
I would say, in general, string matching leaves a lot to be desired, especially when I don't know how to actually type most of the letters of my colleagues's and friends's names. (It's easier on Mac than Windows 10 though - I can't figure out a dang thing for Windows 10 that is not English.)
There is only one key ä on the keyboard, the user probably doesn't have a choice whether that enters the single character version or the two-character version...