I always wondered: if you use reasonable compression (what is reasonably expected to be provided by CDNs these days? brotli? I'm so out of the loop..), if you use this, how much does file size actually matter? Isn't it about entropy, rather than size? Say, if every HTML tag required 136 <'s to open, and 136 > characters to close. How would it actually affect different compression algorithms?
My intuition says that we're all on this wild goose chase for smaller file size while it may (should) not matter at all.
For example, I took Wikipedia's page on entropy and replaced each < by 136*<, same for >. I bring you the file sizes for different algorithms:
normal crazy % larger
raw 423953 3470903 718 %
Gzip -9 78897 134319 70 %
Bzip2 -9 64286 64441 0.24 % (!)
I don't know what compression algo is de rigeur these days, and 136 <s is obviously not the same as substituting double tags for sexprs. Still, I hope we can put this file size boondoggle in the perspective of entropy one day, instead of just mindlessly chasing the character count dragon.EDIT: turns out you can convert HTML to pseudo sexprs with some regex. here are more realistic numbers:
sgml sexpr % diff
raw 423953 403777 -4.8 %
Gzip -9 78897 76936 -2.5 %
Bzip2 -9 64286 63590 -1.1 %
Less radical, but still following the trend.