While I believe we should have a standard, I feel that having UTF-8 in source code brings more risks than benefits. [0]
Next to those additional risks it also limits discoverability if `fiancee` now is written as `fiancée` (yea I had to c/p that from somewhere). Searching for the former does not discover the latter.
Lastly, there is the issue of Intellisense, or whatever it is called in many languages. I have seen codebases in English with the aforementioned `é` in a function name. The only way for me to select that function name was with the arrow keys / mouse. I couldn't type it on my QWERTY.
And yes, I know that there are codebases which are non-English, and not even Latin. Those are very valid concerns to which I admittedly have no answer to.
[0] https://krebsonsecurity.com/2021/11/trojan-source-bug-threat...