To work around this, the international domain name standard defines an encoding called Punycode which maps Unicode to the limited character set DNS supports. The server is unaware of this, and so this optimised tolower() implementation works without any Unicode considerations.
And within unicode, doing it in a "dumb ascii" way probably needs some normalization of diacritics. Eg. 'é' should be U+0065 U+0301 ("e\xcc\x81"), not U+00E9 ("\xc3\x89").
Not sure how punycode handles this, I did once look deeply into it but that was years ago.
[1] https://datatracker.ietf.org/doc/html/rfc7564
[2] https://www.icann.org/resources/pages/root-zone-lgr-2015-06-...
This is the classical label form used, albeit with some additional
restrictions, in hostnames [RFC0952]. Its syntax is identical to
that described as the "preferred name syntax" in Section 3.5 of RFC
1034 [RFC1034] as modified by RFC 1123 [RFC1123]. Briefly, it is a
string consisting of ASCII letters, digits, and the hyphen with the
further restriction that the hyphen cannot appear at the beginning or
end of the string. Like all DNS labels, its total length must not
exceed 63 octets.
I misspoke slightly in my earlier comment: it turns out the DNS protocol does allow octets of any value in a label, but the Internet domain name system -- including all the existing clients, servers, registrars, etc -- as a whole does not. Which is part of why we need a fairly complicated set of specifications to make non-ASCII domain names work.[1] https://datatracker.ietf.org/doc/html/rfc5890#section-2.2