Introducing Brotli: a new compression algorithm for the internet
google-opensource.blogspot.com
google-opensource.blogspot.com
You now have LZ4, Brotil, zstd, snappy, lzfse, lzma all pretty useful practical codec.
Brotil is interesting though. It can be an easy replacement for zlib at level 1 with fairly higher compression then zlib at similar speed.
On the other hand with with lzma it can handily beat lzma but with an even slower compression rate (from the levels that they published in the benchmark) but on the other hand with much higher decompression speed meaning its very good for distribution work. It would be interesting to see the compression ratio, time for levels between 1-9.
It's actually a much easier replacement for zlib then lzma for some. The benchmark shows only level 1,9 and 11. It seems that it can handily beat lzma but at the cost of compression speed (I wonder who use more memory). Then again its decompression speed is so much better making it a perfect choice for distribution.
What truly surprise me though is the work of 'Yann Collet'. A single person so handily beating google's snappy. Even his zstd work looks ground breaking. When I read couple of weeks ago that he was a part time hobby programmer I just didn't know how to be suitably impressed.
Am I reading right that apples lzfse is based of a mix of lz+zstd?
Now we're in a very different place where people upgrade software much faster than in the past and that's likely to start showing up in projects like this where it's actually plausible to think of something like Brotli seeing widespread usage within 1-2 years.
http://fastcompression.blogspot.de/2013/12/finite-state-entr...
http://encode.ru/threads/2313-Brotli?p=44970&viewfull=1#post...
1: https://engineering.linkedin.com/shared-dictionary-compressi...
I prototyped something myself and SHDC for the browser was the closet I could find.
– a more detailed post about Brotli: http://textslashplain.com/2015/09/10/brotli/
– a MSDN article on “compressing the Web”, explaining the difference between deflate and gzip, the limitations of browsers, and mentionning zopfli and brotli: http://blogs.msdn.com/b/ieinternals/archive/2014/10/21/http-...
Seriously, the Brotli quality setting goes to eleven -- pure genius.
http://www.gstatic.com/b/brotlidocs/brotli-2015-09-22.pdf (Fig. 1)
BREACH: Response body compression of a page where there's (a) something attacker controlled, (b) something private and unchanging in the body can reveal that secret, and (c) response length is visible to an attacker. Doesn't require HTTPS.
If an attack applied, it would be one like BREACH. Which isn't surprising: this is a direct replacement for "Accept-Encoding: gzip / Content-Encoding: gzip" and so we should expect it to be in the same security situation.
Even if your page doesn't have sensitive information on it, an insecurely loaded page provides an attacker the avenue to inject potentially malicious code. This will be the case until the entire web is HTTPS-enabled.
They could at least have measured the other algorithms with the same dictionary for fairness.
They won't open source it: the compression tech has for a number of years accounted for a large proportion of Opera's (browser) income. I also don't actually know what OMPD is. Opera Mini server? That's tied sufficiently closely to Presto that it'll only ever get released if Presto does.
The more recent Opera Turbo and the Opera Mini 11 for Android "high-compression mode" (though I have no idea how that really differs to Opera Turbo! [edit: per @brucel it is Opera Turbo]) are certainly available for licensing; Yandex Browser supports Opera Turbo, for example.
Makes sense.
If I saw something like Broetli and didn't immediately recognize the Germanic origin, I could see myself easily misanalyzing it as Bro-etli and pronouncing it that way.
eg: hoe sloe (both rhyme with 'bro')
to get a 'bro-ee' effect you'd need a 'oey' or some other spelling
See: https://en.wikipedia.org/wiki/Germanic_umlaut#Orthography_an...
Also, simply dropping those dots where they actually belong is refreshingly post-metal-umlaut.
https://d3ijcvgxwtkjmf.cloudfront.net/a4c3c7313b7bdeb68ad46a...
Base: 6,779,000 bytes
GZip2 Normal: 2,296,362 bytes, Ultra: 2,258,967 bytes
LAMZ2 Normal: 921,600 bytes, Ultra: 920,147 bytes
Zstd is better for dynamic compression. Or libslz today for very fast compression with zlib compatibility.
I even do it with my name most of the times. A Chinese official that sees Ø in my passport won't understand why I write 'OE', and might even start questioning if it is the same name, but 'O' never fails.
Airlines can read the gibberish in the bottom of my passport and see that the transliteration is actually 'OE', but most places don't have a scanner for that.
For comparison here's a chart of hex character and approximate [2] number of occurrences.
(0 - 15k) (1 - 10k) (2 - 19k) (3 - 14k) (4 - 15k) (5 - 16k) (6 - 62k) (7 - 32k) (8 - 10k) (9 - 11k) (a - 11k) (b - 7k) (c - 8k) (d - 12k) (e - 20k) (f - 9k)
[1]: http://www.ietf.org/id/draft-alakuijala-brotli-05.txt [2]: counted with find-in-page, didn't bother to only search the dict
Reading the PDF paper I see they use for benchmarking Xeon with price-tag in range of $560, how this can be reconciled with hot phrases like boosting browsing on mobile devices?!
"Input: 812,392,384 bytes, HTML top 10k Alexa crawled (8,998 with HTML response)
Output: 219,591,148 bytes, 1.799 sec., 1.355 sec., tor-small -1
218,018,012 bytes, 1.951 sec., 1.464 sec., tor-small -1 -b800mb
210,647,258 bytes, 1.736 sec., 1.328 sec., qpress64 -L1T1
210,194,996 bytes, 0.333 sec., 0.371 sec., lz4 -1
194,233,793 bytes, 2.818 sec., 1.348 sec., qpress64 -L2T1
187,766,706 bytes, 7.966 sec., 1.059 sec., qpress64 -L3T1
173,904,470 bytes, 2.995 sec., 2.721 sec., tor-small -2
173,418,150 bytes, 3.132 sec., 2.843 sec., tor-small -2 -b800mb
169,476,113 bytes, 2.072 sec., 0.352 sec., lz4 -9
165,571,040 bytes, 2.931 sec., 2.820 sec., NanoZip - f
158,213,503 bytes, 1.855 sec., 0.980 sec., zstd
154,673,082 bytes, 3.213 sec., 2.445 sec., tor-small -3
154,166,902 bytes, 3.364 sec., 2.555 sec., tor-small -3 -b800mb
152,477,067 bytes, 2.128 sec., 0.973 sec., lzturbo -30 -p1
152,477,067 bytes, 2.132 sec., 0.971 sec., lzturbo -30 -p1 -b4
150,773,269 bytes, 2.151 sec., 1.047 sec., lzturbo -30 -p1 -b16
150,219,553 bytes, 2.332 sec., 1.204 sec., lzturbo -30 -p1 -b800
149,670,044 bytes, 7.825 sec., 2.567 sec., WinRAR - 1
149,642,742 bytes, 2.069 sec., 0.646 sec., zhuff_beta -c0 -t1 145,770,266 bytes, 4.586 sec., 2.285 sec., tor-small -4
141,951,602 bytes, 4.484 sec., 2.360 sec., tor-small -4 -b800mb
141,215,050 bytes, 2,751 sec., 0.938 sec., lzturbo -31 -p1 -b4
140,657,806 bytes, 5.037 sec., 2.544 sec., FreeArc - 1
140,211,060 bytes, 6,970 sec., 1.775 sec., bro -q 1
138,103,483 bytes, 118.023 sec., 2.394 sec., cabarc -m LZX:15
138,051,401 bytes, 2.761 sec., 1.001 sec., lzturbo -31 -p1 -b16
137,564,310 bytes, 18.762 sec., 4.299 sec., NanoZip - dp
137,211,547 bytes, 7.808 sec., 1.712 sec., bro -q 2
137,000,208 bytes, 3.763 sec., 3.852 sec., NanoZip - F
136,523,335 bytes, 2.830 sec., 1.094 sec., lzturbo -31 -p1
136,445,854 bytes, 50.932 sec., 2.344 sec., lzhamtest_x64 -m0 -d24 -t0 -b
136,337,495 bytes, 14.823 sec., 4.259 sec., NanoZip - d
135,723,691 bytes, 8.318 sec., 1.677 sec., bro -q 3
135,656,476 bytes, 2.972 sec., 1.153 sec., lzturbo -31 -p1 -b800
135,315,388 bytes, 51.436 sec., 2.371 sec., lzhamtest_x64 -m0 -t0 -b
135,287,357 bytes, 51.937 sec., 2.418 sec., lzhamtest_x64 -m0 -d29 -t0 -b
135,071,576 bytes, 21.650 sec., 4.251 sec., NanoZip - dP
132,819,515 bytes, 20.102 sec., 6.278 sec., 7-Zip - 1
131,871,664 bytes, 14.052 sec., 0.899 sec., lzturbo -32 -p1 -b4
131,401,865 bytes, 9,917 sec., 1.677 sec., bro -q 4
129,184,341 bytes, 8.305 sec., 2.692 sec., tor-small -5
127,355,215 bytes, 20.866 sec., 5.825 sec., 7-Zip - 2
127,045,472 bytes, 9.549 sec., 0.957 sec., lzturbo -32 -p1 -b16
126,139,033 bytes, 8.025 sec., 2.751 sec., tor-small -5 -b800mb
125,732,647 bytes, 10.642 sec., 2.618 sec., tor-small -6
125,454,769 bytes, 140.513 sec., 2.281 sec., cabarc -m LZX:18
123,169,077 bytes, 8.472 sec., 1.090 sec., lzturbo -32 -p1
123,093,411 bytes, 22.468 sec., 5.508 sec., 7-Zip - 3
122,564,329 bytes, 10.074 sec., 2.680 sec., tor-small -6 -b800mb
122,480,456 bytes, 19.411 sec., 1.645 sec., bro -q 5
121,068,548 bytes, 14.536 sec., 3.289 sec., FreeArc - 2
120,653,755 bytes, 16.107 sec., 2.552 sec., tor-small -7
119,969,489 bytes, 27.663 sec., 1.602 sec., bro -q 6
119,740,393 bytes, 27.370 sec., 5.259 sec., 7-Zip - 4
118,343,545 bytes, 24.112 sec., 2.123 sec., WinRAR - 2
118,139,032 bytes, 35.361 sec., 4.371 sec., NanoZip - Dp
117,500,327 bytes, 15.517 sec., 2.594 sec., tor-small -7 -b800mb
117,388,039 bytes, 23.546 sec., 2.524 sec., tor-small -8
116,526,595 bytes, 37.847 sec., 4.383 sec., NanoZip - DP
116,454,906 bytes, 35.232 sec., 4.351 sec., NanoZip - D
116,269,246 bytes, 25.888 sec., 6.589 sec., FreeArc - 3
116,217,001 bytes, 40.748 sec., 1.630 sec., bro -q 7
115,993,125 bytes, 192.929 sec., 2.199 sec., cabarc -m LZX:21
115,985,847 bytes, 386.192 sec., 0.850 sec., lzturbo -39 -p1 -b4
115,729,606 bytes, 35.504 sec., 2.095 sec., WinRAR - 3
115,163,486 bytes, 55.523 sec., 1.614 sec., bro -q 8
115,022,074 bytes, 49.863 sec., 2.084 sec., WinRAR - 5
114,602,026 bytes, 8.403 sec., 1.218 sec., lzturbo -32 -p1 -b800
114,345,025 bytes, 78,418 sec., 1.594 sec., bro -q 9
114,281,170 bytes, 22.925 sec., 2.575 sec., tor-small -8 -b800mb
113,354,128 bytes, 29.519 sec., 2.474 sec., tor-small -9
112,376,531 bytes, 177.077 sec., 1.923 sec., lzhamtest_x64 -m1 -d24 -t0 -b
111,848,802 bytes, 29.046 sec., 2.515 sec., tor-small -9 -b800mb
110,532,234 bytes, 40.580 sec., 2.496 sec., tor-small -10
110,177,215 bytes, 54.632 sec., 6.398 sec., FreeArc - 4
109,908,468 bytes, 40.292 sec., 2.501 sec., tor-small -10 -b800mb
109,522,695 bytes, 208.748 sec., 1.898 sec., lzhamtest_x64 -m2 -d24 -t0 -b
109,425,530 bytes, 436.824 sec., 0.893 sec., lzturbo -39 -p1 -b16
108,520,934 bytes, 58.227 sec., 2.521 sec., tor-small -11
108,520,934 bytes, 58.329 sec., 2.518 sec., tor-small -11 -b800mb
107,850,398 bytes, 266.166 sec., 2.562 sec., tor-small -12
107,842,909 bytes, 267.559 sec., 2.550 sec., tor-small -12 -b800mb
106,128,420 bytes, 271.607 sec., 1.850 sec., lzhamtest_x64 -m3 -d24 -t0 -b
105,933,030 bytes, 571,280 sec., 5.168 sec., lzturbo -49 -p1 -b4
105,692,791 bytes, 193.962 sec., 1.919 sec., lzhamtest_x64 -m1 -t0 -b
104,539,771 bytes, 316.307 sec., 1.833 sec., lzhamtest_x64 -m4 -d24 -t0 -b
104,094,380 bytes, 2313.780 sec., 1.693 sec., bro -q 10
104,053,219 bytes, 195.503 sec., 1.977 sec., lzhamtest_x64 -m1 -d29 -t0 -b
102,895,078 bytes, 148.997 sec., 4.850 sec., 7-Zip - 5
101,941,653 bytes, 237.364 sec., 1.889 sec., lzhamtest_x64 -m2 -t0 -b
100,898,120 bytes, 534.627 sec., 1.097 sec., lzturbo -39 -p1
100,159,922 bytes, 239.813 sec., 1.933 sec., lzhamtest_x64 -m2 -d29 -t0 -b
99,699,129 bytes, 625.001 sec., 4.902 sec., lzturbo -49 -p1 -b16
96,239,572 bytes, 347.011 sec., 1.893 sec., lzhamtest_x64 -m3 -t0 -b
95,197,295 bytes, 236.139 sec., 4.587 sec., 7-Zip - 9
94,133,011 bytes, 356.440 sec., 1.933 sec., lzhamtest_x64 -m3 -d29 -t0 -b
93,601,386 bytes, 431.884 sec., 1.899 sec., lzhamtest_x64 -m4 -t0 -b
92,303,359 bytes, 727.475 sec., 4.729 sec., lzturbo -49 -p1
91,310,894 bytes, 449.102 sec., 1.923 sec., lzhamtest_x64 -m4 -d29 -t0 -b
90,239,627 bytes, 680.976 sec., 1.170 sec., lzturbo -39 -p1 -b800
87,715,022 bytes, 314.169 sec., 5.428 sec., FreeArc - 9
82,891,405 bytes, 882.513 sec., 4.597 sec., lzturbo -49 -p1 -b800
77,286,010 bytes, 6497.059 sec., 7.715 sec., glza
Used: 7z 15.07 beta - Sep 17, 2015 (one thread)
rar 5.40 beta 4 - Sep 21, 2015 (one thread)
arc 0.67 - Mar 15, 2014 (one thread)
nz 0.09 - Nov 4, 2011 (one thread)
zhuff_beta 0.99 - Aug 11, 2014
cabarc 6.2.9200.16521 - Feb 23, 2013
lz4 1.4 - Sep 17, 2013
qpress64 1.1 - Sep 23, 2010
zstd 0.0.1 - Jan 25, 2015
tor-small 0.4a - Jun 2, 2008
lzturbo 1.2 - Aug 11, 2014
lzhamtest_x64 1.x dev - Sept 25, 2015 (own VS2015 compile)
glza 0.3a - Jul 15, 2015 "
lzturbo seems essentially better, a few other happen to be better.
In my view the two best are:
169,476,113 bytes, 2.072 sec., 0.352 sec., lz4 -9
115,985,847 bytes, 386.192 sec., 0.850 sec., lzturbo -39 -p1 -b4
LzTurbo benefits a lot from bigger blocks, in above example it uses 4MB block single-threadedly, and look how fast it is, if 16 threads are to be used what the outcome would be ... maybe 0.100 sec?!
fastest encoding:
210,194,996 bytes, 0.333 sec., 0.371 sec., lz4 -1
fastest decoding:
169,476,113 bytes, 2.072 sec., 0.352 sec., lz4 -9
149,642,742 bytes, 2.069 sec., 0.646 sec., zhuff_beta -c0 -t1
gzip-alternative:
135,723,691 bytes, 8.318 sec., 1.677 sec., bro -q 3
114,602,026 bytes, 8.403 sec., 1.218 sec., lzturbo -32 -p1 -b800
prepacked:
104,094,380 bytes, 2313.780 sec., 1.693 sec., bro -q 10
90,239,627 bytes, 680.976 sec., 1.170 sec., lzturbo -39 -p1 -b800
best compression:
82,891,405 bytes, 882.513 sec., 4.597 sec., lzturbo -49 -p1 -b800
77,286,010 bytes, 6497.059 sec., 7.715 sec., glza
there are missing e.g. density, zpaq, and comments there suggest that brotli doesn't look that good for other than text data ...
As one (Jyrki) of co-authors commented: ``` For more clarity on the situation, you could compare LZMA, LZHAM and brotli at the same decoding memory use. Possibly values between 1-4 MB (window size 20-22) are the most relevant for the HTTP content encoding. Unlike in a compression benchmark, there are a lot of other things going on in a browser, and the allocated memory at decoding time is a critically scarce resource. ``` Source: http://google-opensource.blogspot.bg/2015/09/introducing-bro...
- 10Mbps or 1MB/s connection;
- 100Mbps or 10MB/s connection.
The goal is to receive in our browser those 812,392,384 bytes as quickly as possible, in first case the winner is 77,286,010 bytes compressor, in second the winner is 90,239,627 bytes compressor, yes?
In first case transfer_time + decompression_time = 77s + 7s = 84s
In second case transfer_time + decompression_time = 9s + 1s = 10s
Now, you see that even in Web Browsing scenario the best performer is not established, right?
html8 : 100MB random html pages from a 2GB Alexa Top sites corpus. number of pages = 1178 average length = 84886 bytes.
The pages (length + content) are concatenated into a single html8 file, but compressed/decompressed separately. This avoids the cache scenario like in other benchmarks, where small files are processed repeatedly in the L1/L2 cache, showing unrealistic results.
size: 100,000,000 bytes. Single thread in memory benchmark cpu: Sandy bridge i7-2600k at 4.2 Ghz, all with gcc 5.1, ubuntu 15.04
size ratio% C MB/s D MB/s MB=1.000.000
15180334 15.2 0.43 482.07 brotli 11 v0.2.0
15309122 15.3 2.27 127.23 lzma 9 v15.08
16541706 16.5 2.07 1463.39 lzturbo 39 v1.3
16921859 16.9 2.96 230.54 lzham 4 v1.0
17153795 17.2 0.13 474.63 zopfli v15-05
17860382 17.9 43.51 495.78 zlib 9 v1.2.8
18033576 18.0 135.62 1454.31 lzturbo 32 v1.3
100000000 100.0 5984.00 6043.00 libc memcpy
LzTurbo compress 5 times and decompress 3 times faster than brotli.LzTurbo decompress more than 6 times faster than lzham
I wonder if anyone working on the field would have a comment on that.
make
Worked for me on Ubuntu 14.04. Not sure what packages are necessary. YMMV.
$ cd brotli
python extension:
$ python setup.py build
static lib:
$ cd enc
$ make
$ ar rvs brotli.a *.o
Strangely enough, no one compressor is better in all situations. :)
https://github.com/cloudflare/zlib
Previously discussed here:
https://news.ycombinator.com/item?id=9857784
And a recent performance comparison with Intel's patches:
https://www.snellman.net/blog/archive/2014-08-04-comparison-...
Says the company whose business model consists of convincing publisher to put ads in their website, thus slowing down page load times. Same company that lets you integrate Web elements like scripts and web fonts from their servers, once again making your webpage slower to load.
I know I'm exaggerating it a bit, but I really hate this company mission pitches that people in big companies constantly use to open their technical communications.
For me, it's akin to state propaganda in non-democratic countries: everyone knows that whatever authorities say and what the truth is are two different things, so there's little point to even analysing the official message (except maybe for humor factor etc.)
If they thought my time was valuable, they would not put colored strips at the top of every screen that say "Switch to chrome" or "Switch to GMail" or even after I've switched to Chrome, "Chrome is not your default browser." If they thought my time was valuable, they wouldn't be wasting it with their constant desperate pleas to browse the internet in precisely the way they want. If they really valued my time, they wouldn't throw random context switches into every single one of their web properties for whatever is the product du jour.
If they wanted my page to load faster, they would work as hard on making their stupid website addons load fast as they worked on getting everyone on earth to install them. Running local mirrors of ajax.googleapis.com and fonts.googleapis.com, along with hijacking analytics and doubleclick and returning 0-byte files, is the best thing I ever did for my poor parents' satellite internetion connection.
Internet users' time (and site loading speeds) are very much second- or third-class items on google's list of things to give a shit about. I'm absolutely not questioning thees decisions; google is an advertising agency and must prioritize this over my convenience! But pretending like they're some kind of altruistic charity working for the common good is disingenuous and slightly offensive.
It also creates the problem where some of the more gullible people actually believe google, a huge faceless global company, gives a shit about anyone in particular, which is demonstrably untrue.
Does the fact that Walmart charges money for their products make the above sentence disingenuous? Of course Walmart charges money, that's how their company works. That doesn't mean they can't try to reduce waste or inefficiency in their system.
The acts of a person implies that person's motivation is not to use a product or comment about qualities a product but to bash the producer/creator using the context of the product. So ulterior motive is to punish creator, the actual product is only a medium. (grape: product - wineyard keeper: producer)
Sorry it sounded a bit harsh. I also agree that a dry and more technical blog entry would be better, but I think it is not a big deal considering the importance of the product.
There is plenty of evidence of significant drop in users numbers with increasing page load time.
An advertising company has to ensure that their ads are not only displayed but that there is also a reasonable chance that they will be noticed. Otherwise it becomes a pointless exercise, advertisers leave and revenues drop.
Therefore what is really going on here is Google trying to squeeze in more advertising before annoying the users to the point of leaving the page.
Yes, I agree, advertising in itself is inimical to user's time and thus the above sentence can be seen as deeply hypocritical.
;)
Is this the same story everytime that G. invents something, implements it in Chrome and then gains advantage, because takes it as standard and other browser hadn't implement it yet?
This wasn't developed and deployed in secret. According to the comments in https://bugzilla.mozilla.org/show_bug.cgi?id=366559, the GitHub repository has been public since at least November 2014.
You may not like Google, but insinuating evil plans every time they do something cool isn't helping anyone.
By the way, new features are generally created this way. They are added to browsers way before they are standardized. You see, convincing the other browser vendors isn't easy. You need very compelling arguments.
Take WebP, for example. It provides massive benefits. Especially if you can use lossy RGBA instead of PNG32 (easily 80% smaller). And yet, Mozilla shows little interest in implementing it.