You are right that for the majority of applications json is fine.
You are right that for the majority of applications json is fine.
Isn't that the whole point of compression: to minimise repetition of common features in data?
ASCII is extremely repetitive. 1MB of random braces gzips down to 163KB. That's a compression ratio of 6.1:1. Optimal compression would be close to 8:1 (since each byte really contains only one bit of information).
Your sample data is bunk for determining the compression characteristics of braces in a realistic json file.
Huffman coding works by replacing common symbols with shorter codes and uncommon symbols with longer codes.
The codes used are therefore variable-length.
The codes are created in such a way that no code is the prefix of any other code.
That way, the decoder is able to know when it has reached the end of a code without the need for extra information other than the Huffman tree, which tells the decoder which codes belong to which symbols.
The Huffman tree can be pre-agreed or, more commonly, included with the compressed data.