It's not like internet standards don't know there are other languages, it's just that they documented how things were done at the time. Some legacy has remained ever since.
Unicode-aware strings are the right choice for 99% of code. The last 1% should be a special case.
You don't need a raw socket to get into trouble. You also don't need a "legacy" protocol.
I’d wager most code is application code, where UTF8 strings are a great choice.
IIRC historically on Windows, a string was UTF16, on unix it was ASCII; nowadays everywhere it's UTF8 without a way to specifically limit what goes into a string.
For example UTF8 opens the door to homoglyph attacks and various other things (RTL, spaces), and a program should be able to force a string to be ASCII-only so that these classes of problems are ruled out.
That being said, Windows permits unpaired surrogates in its "UTF-16" strings, even though that is not actually valid UTF-16. Similarly, many Linux APIs accept arbitrary bytes, not just valid UTF-8.