CDN77 Now Supports Brotli
blog.cdn77.com
blog.cdn77.com
https://blog.cloudflare.com/results-experimenting-brotli/
If you want to experiment with it, it's enabled on CloudFlare's test server https://http2.cloudflare.com/.
Since we don't charge by byte the focus is on end-user performance and not 'saving money' and we concluded that it wasn't currently worth implementing Brotli widely but will continue to experiment with it.
"Most files are smaller than 64KB, and if we look only at those files then Brotli 4 is actually 1.48X slower than zlib level 8!"
And the faster zlib levels (1-7) are even faster! See the table. I especially like zlib 1.
All these can't be compressed more than they already are, forcing any lossless (like zlib and brotli) compressor on them is just the loss of CPU time.
Try and compare.
Compression 101.
$ curl -sS http://www.exiv2.org/include/img_1771.jpg | wc -c
32764
$ curl -sS http://www.exiv2.org/include/img_1771.jpg | gzip -9 | wc -c
31323
In this case, though, it often makes more sense to just remove the EXIF data.PDF files can be even more compressible, at least ones that are text and layout as opposed to embedded images:
$ curl -sS http://www.polyu.edu.hk/iaee/files/pdf-sample.pdf | wc -c
7945
$ curl -sS http://www.polyu.edu.hk/iaee/files/pdf-sample.pdf | gzip -9 | wc -c
4336Those that want to deliver the videos or pictures should use the proper format and not depend on Brotli or zlib.
CloudFlare is very close to end users (single digit milliseconds) and gzip is fast and gives good compression. What we can't afford to do is make a file smaller at the cost of increased latency because the compression time was longer. There's a trade off between compression time and delivery time.
So, Brotli makes sense for things we cache (because we can compress out of band) and for large files delivered over slow connections but is not a panacea.
I work a block from your office in SF; may I make a request: can you guys host a public tech-talk so I can come to your office and hear it straight from your engineers and maybe have the opportunity to ask questions.
Your Writeup about the HFT NICs from solar(something) circa ~2012 has had me hooked on cloudflares opinion on pretty much everything.
The only other Corp blog post I love as much as yours are the ones from Backblaze re their pods.
It's at least a very good codec, though, so it's still a win for other data. Just smaller than you might expect.
Edit: They wrote this about it in http://www.gstatic.com/b/brotlidocs/brotli-2015-09-22.pdf :
> Unlike other algorithms compared here, brotli includes a static dictionary. It contains 13’504 words or syllables of English, Spanish, Chinese, Hindi, Russian and Arabic, as well as common phrases used in machine readable languages, particularly HTML and JavaScript. The total size of the static dictionary is 122’784 bytes. The static dictionary is extended by a mechanism of transforms that slightly change the words in the dictionary. A total of 1’633’984 sequences, although not all of them unique, can be constructed by using the 121 transforms. To reduce the amount of bias the static dictionary gives to the results, we used a multilingual web corpus of 93 different languages where only 122 of the 1285 documents (9.5 %) are in languages supported by our static dictionary.
https://gist.github.com/xnyhps/677f7c1b444f346bef99
(I cleaned it up a bit to remove newlines and tabs, and a couple that are entirely of unprintable characters.)
Brotli is a great compressor, especially at levels 2-5. Unfortunately, the Google paper on Brotli runs tests at levels 1 and 11. I don't get that at all when their stated goal was to replace gzip.
See static Huffman coding in fax machines.
The claims that it is comparable to xz/lzma for generic binary data are not accurate.
In my real-world tests of compressing 3D data it far underperformed xz/lzma although it was still better than gzip:
You will get much better results out of Brotli if you restructure your data to be more compressible, and that will also improve your lzma and gzip (especially gzip) compression ratios, to a tremendous degree. Have you done any of this? If not, ping me, and I can explain some techniques to apply.
This sounds interesting, I'd like to read some examples/links/explanations.
It'd be an interesting test to take our eight-language set of localization strings and compress them in UI order and language order and see if there's much of a difference. (UI order is all languages for one dialog element, then all eight for the next, etc. Language order is all the English first, then ...)
I'd definitely like to hear Kevingadd's tips though.
Similarly, an entire IOS device could be fabricated as a single ASIC, and (for example) uikit could be fabricated as part of that ASIC.
There is always a tradeoff between generic optimization and usage-specific optimization which comes at the expense of flexibility.
Google can do a statistical analysis of all the data it serves compressed with gzip, and determine exactly the characteristics of a compression algorithm that would save the most money.
These are small, evolutionary optimizations that save tons of money by incrementally increasing efficiency in a large system.
"In late September, Google released a compression algorithm called Brotli and gave files it makes the extension “.bro”.
But last week the extension was changed to “.br”.
The reason for the change is threads like this one, in which posters suggest that “'bro' has a gender problem” and “comes of[f] misogynistic and unprofessional due to the world it lives in.”
http://www.theregister.co.uk/2015/10/11/googles_bro_file_for...
modern filesystems can handle a few extra characters
.br<tab> is the same number of characters as .bro
So, instead of using a word which is mostly used as a friendly term of familiarity/endearment, they decided to go with a word which has connotations of racism.
I... can't even.
It's the world we live in today and much better to just avoid altogether now rather than try to defend/fix it later.
Now, you can not like that definition of bro, but that is for many people to common definition. The first thing that comes to mind when I hear 'bro' is the phrase 'bros before hoes', which is hard to interpret as not offensive to women.
Does this really connect to a compression format? Of course not, but if that is the first thing that comes to people's minds, why not switch to something with less baggage?
If you can't see how any of this could be offensive, then I'm afraid we just have different points of view.
Everything can be offensive to someone if they try hard enough. Just pick something that doesn't intend to be rude and ignore the few haters who have to make everything about them.
This isn't the most serious thing in the world, but there are definitely people who find it excluding, and that seems reasonable to me.
This is CS, a lack of groupthink-y "professionalism" is why I like this field.
While it's not nice to accept, most women I know well in computing have had bad experiences with "bro"-type people, saying they shouldn't / can't use a computer because they are a woman. On an at least monthly basis. For years. It grinds slowly over time.
If you think CS is professional, I've got some bad news for you -- there are quite a lot of toxic badly behaved people around unfortunately.
But my point is that this is not particularly offensive, it's purely a pun. I think being oversensitive can be as damaging as being unsensitive.
But I disagree with how hard you have to try. The people on here are trying fairly hard to make sure everyone knows how offensive they might find it. It didn't offend you, but you're offended that it might offend someone, or worse - that it might not offend someone. Nobody said "It's the compression method for white men", that whole racist angle is yours.
I find your shaming word-police game to be exclusionary. Please stop it.
Who exactly is being excluded? No-one (it seems to me) watch attached to .bro, it only existed for a couple of days.
If anyone did like it I doubt they'd have felt free to speak up.
> I do know people who found it offensive
I find Bros offensive - in my house. But the word? No. "Nazi" is just a word, Nazis are offensive.
> I can see no particular reason not to change the name.
Ditto there's no reason to. Someone went out of their way to take offense which was bounced around an echo chamber and now everyone is offended by something they didn't know existed before today.
That's not something we should reward.
I did look around for known brotli+Firefox+nginx issues, but the only one that comes up is https://bugzilla.mozilla.org/show_bug.cgi?id=1215724 which was fixed in shipping Firefox months ago, so I'm assuming that's not the one you're talking about.
https://hg.mozilla.org/releases/mozilla-aurora/log/10e1774de...
Which commit are you talking about, exactly? Bug 1242904 was backported to Firefox 45, and everything else I see at http://hg.mozilla.org/mozilla-central/filelog/tip/modules/br... as of today (which is the same as <http://hg.mozilla.org/mozilla-central/filelog/ea6298e1b4f7/m...) was checked in way before 45....
Of course the version update wasn't supposed to change the on-the-wire behavior... or so the library authors claimed. :(
https://hacks.mozilla.org/2015/11/better-than-gzip-compressi...
https://github.com/ende76/brotli-rs
It's currently in use in Servo
I also don't understand why not xz, though I know that requires significantly more resource than gzip.
If you're saying brotli compresses/decompresses faster that gzip, it doesn't according to Cloudflare tests [1] (see table near bottom).
Even the fastest compression level of Brotli is slower than the highest/slowest compression level of gzip in most cases.
[1]: https://blog.cloudflare.com/results-experimenting-brotli/
It doesn't always compress faster, but as far as I can tell they didn't measure or say anything about decompression.
time bro --quality 6 --input linux-4.5.tar --output linux-4.5.tar.bro
real 0m18.436s user 0m18.228s sys 0m0.184s
time gzip -9 linux-4.5.tar real 0m30.555s user 0m30.424s sys 0m0.172s
ls -lh (I removed the metadata): 106M linux-4.5.tar.bro 129M linux-4.5.tar.gz
Only 82% of the size at 60% of the time. Now this is pure text so that's not a good example of everything.
re: "Brotli should bring 25% reduction in data size compared to Gzip for the most common assets like Javascript and CSS files. For HTML, Brotli promises up to 40% difference (with median around 25%)."
https://cran.r-project.org/web/packages/brotli/vignettes/bro...
Not sarcasm, real question.