Brotli Compressed Data Format
tools.ietf.org
tools.ietf.org
For example a paper from last SODA [1] shows that by using a better optimization algorithm and just simple universal codes instead of Huffman, it is possible to beat zlib in space and compete with Snappy in decompression speed (and I'm sure that the decoder can be further optimized). This is just the last one on the topic but there have been quite a few.
Caveats: compression is fairly slow (but can be improved with heuristics) and the datasets used in the experiments are big and very repetitive (so it is not clear how it performs on small files), but still I think that many ideas should finally find their way into modern compression formats.
[1] http://arxiv.org/pdf/1307.3872v1.pdf
Disclaimer: the paper authors are from my CS department.
Brotli has something else in common with Snappy: a lack of a specified framing mechanism. When Snappy first came out it did not have one, so if you wanted to write Snappy-compressed data to a file, you might invent your own header to frame the compressed data. They later added this to Snappy, but by then it was too late: implementations had already come up with their own, mutually-incompatible framing. The same mistake seems to be being made again with Brotli.
onclick="javascript:
was common enough in their corpus to end up in the dictionary.Readable version: https://gist.github.com/anonymous/f66f6206afe40bea1f06
onclick="http:
and it would work, for the same reason. 12010. "University of"
12129. "University of "
13106. "the University of"
13227. "Oxford University"
13534. " Oxford University"
(I'm sure there's other similar duplicates; I just happened to notice them while looking at all the "university"-ies in the corpus.)Especially if we know it's supposed to be used to compress the fonts.
Incidentally I really like LZ77 and its variants; it's my favourite compression algorithm, due to its incredible simplicity (a decompressor is just a bit over a dozen instructions) and intuitiveness. Thus I've always considered it odd that Huffman's paper was more than 20 years before LZ's - the Huffman algorithm is rather more complex. Perhaps the LZ algorithm was well known already, and considered too trivial to write a paper about? I do know that many have rediscovered LZ independently, without ever studying data compression theory.
Can be produced or consumed, even for an arbitrarily long sequentially presented input data stream, using only an a priori bounded amount of intermediate storage, and hence can be used in data communications or similar structures, such as Unix filters;
It's supposed to be a safe compression standard.
https://code.google.com/p/font-compression-reference/source/...
http://lists.w3.org/Archives/Public/public-webfonts-wg/2013O...
http://lists.w3.org/Archives/Public/public-webfonts-wg/2013O...
EDIT: They are from October 2013, so they may be inaccurate.
WOFF report on Brotli: http://www.w3.org/TR/WOFF20ER/#candidateb (and #candidatea says what they thought of LZMA)
Reference code: https://code.google.com/p/font-compression-reference/ (this aims for very slow but good compression, like their zopfli zlib encoder)
Google's presentation: https://docs.google.com/presentation/d/1aigINmRR7fw_ml8rz0rJ...
To the authors, I suspect it was a feature that it's 'just' DEFLATE with 4MB windows + more context for entropy encoding + tuned-up coding + a static dictionary + some other stuff--the WOFF spec mentions similarty to DEFLATE as a benefit, and they would like to keep it easy enough to for others to write decoders.
I kind of wonder if it was derived from something Google had been using internally. If you decode the static dictionary in the spec, it has a bunch of words and phrases from the most-used languages as you'd expect, but also some things that look like common code fragments from Web pages, which would be right down Google's alley. And that's not stuff you need for font compression in particular.
https://docs.google.com/presentation/d/1aigINmRR7fw_ml8rz0rJ...
[1] https://code.google.com/p/gipfeli/ [2] https://code.google.com/p/zopfli/