> These files are sometimes NOT using the ASCII encoding but UTF-8, e.g. Java source files.
Why is this a problem? In 99.9999% of Java code, the UTF-8 characters aren't going to trip up an ASCII search.
Why is this a problem? In 99.9999% of Java code, the UTF-8 characters aren't going to trip up an ASCII search.
That doesn't even display on my browser[1]; tried it in Goland[2], doesn't display there either, so that's the rare case 0.0001% that I wouldn't really worry about, because if the code has undisplayable unicode sequences, there's bigger problems than searching.
[1] Chrome, on Mac
[2] Also on Mac
To demonstrate this OP explicitly used "DOTTED CIRCLE" (◌) then added the "COMBINING ACUTE ACCENT" to that. Normally there would be no dotted circle.
I wrote an article on this a while ago: https://richardjharris.github.io/unicode-in-five-minutes/