https://unicode.org/reports/tr10
tl;dr from its wikipedia page:
> The Unicode collation algorithm (UCA) is an algorithm defined in Unicode Technical Report #10, which is a customizable method to produce binary keys from strings representing text in any writing system and language that can be represented with Unicode. These keys can then be efficiently compared byte by byte in order to collate or sort them according to the rules of the language, with options for ignoring case, accents, etc
A
B
% printf "%s" A B | xxd -b -c4
00000000: 11110000 10011101 10010000 10110100 ....
00000004: 11110000 10011101 10010000 10110101 ....
% printf "%s" A B | xxd -c1 -ps | sort | xxd -r -ps | xxd -b -c4
00000000: 10010000 10010000 10011101 10011101 ....
00000004: 10110100 10110101 11110000 11110000 ....
% printf "%s" A B | xxd -c1 -ps | sort | xxd -r -ps
????????