OP converted a Unicode file to an ascii file.
Isn’t the loss of precision what is “compressing” the file? The Shakespeare corpus probably doesn’t use any Unicode characters.
Isn’t the loss of precision what is “compressing” the file? The Shakespeare corpus probably doesn’t use any Unicode characters.
An encoding (like ASCII or UTF-8) describes how bytes are mapped to such symbols. ASCII only describes 128 characters whereas UTF-8 can represent any Unicode symbol in bytes.
Notably though, UTF-8 is a strict superset of ASCII. That is, every character ASCII represents is represented by the same bytes in UTF-8. So encoding Romeo and Juliet in ASCII or UTF-8 results in the same exact file.