I wish nothing was in UTF-8 and UTF-8 was relegated to properties files. There are codebases out there with complete i18n and l10n in more languages that most here have ever worked with where there's zero Unicode characters allowed in source code files (with pre-commit hooks preventing committing such source code files).
Bruce Schneier was right all along in 1998 or whatever the date was when he said: "Unicode is too complex to ever be secure".
We've seen countless exploits based on Unicode. The latest (re)posted here on HN was a few days ago: some Unicode parsing but affecting OpenSSL. Why? To allow support for internationalized domain names and/or internationalized emails.
Something that should never have been authorized.
We don't need more of what brings countless security exploits: we need less of it.
Relegated Unicode to translation/properties file, where it belongs.
Sure, Unicode is great for documents, chat, etc.
But everything in UTF-8? emails? domain names? source code? This is madness.
I don't understand how anyone can admire the fact that HANGUL fillers are valid in source code are somehow a great win for our industry.