That's interesting. Is that a common optimization? I hadn't heard of any other web server doing that.
That's interesting. Is that a common optimization? I hadn't heard of any other web server doing that.
It's not terribly exact, +-1000 bytes would probably not make a big difference, but I think it's a good default.
And of course, some people may have unique use cases where a custom threshold may be better.
gzip on;
gunzip on;
gzip_http_version 1.1;
gzip_proxied any;
gzip_comp_level 5;
gzip_disable "msie6";
gzip_vary on;
gzip_min_length 2048;If an attacker controls /some/ of this data, and would like to read /other parts/, they can abuse compression to measure whether the parts they don't know are "like" the part they control, because if they are then the compression will make the results shorter than otherwise which they can passively measure.
It's not a problem to move a compressed object over a secure channel on its own, the problem arises if either you try to compress the channel which is moving objects from different origins (e.g. a cookie set by a random advertising web site and your Facebook password) or compress a composite object e.g. maybe your backups mixed with a file you downloaded from a dodgy "pirate" video site.
This was mentioned in 2012 when CRIME[1] (also included BEAST[2] exploit) and later BREACH[2] vulnerability (when it was considered cool to come up with cool sounding name, creating a logo and a website for specific vulnerabilities)
[1] https://en.wikipedia.org/wiki/CRIME
[2] https://en.wikipedia.org/wiki/Transport_Layer_Security#BEAST...
For example, https://github.com/expressjs/compression/blob/dd5055dc92fdea...
> Even if both the client and the server supports the same compression algorithms, the server may choose not to compress the body of a response, if the identity value is also acceptable.
I've actually run into this twice in my career and it has been a surprise to those around me in both cases. Both times in the context of small payloads where the server is applying some heuristic about whether to encode or not. (e.g status page stops sending gzipped output when the server is becoming "unhealthy")
[0]: https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Ac...
This makes obvious sense once you consider that the client tells the server which compression formats it supports in every request yet not every data format is compressible, nor does the server necessary support any candidate compression format.
For example, the server wouldn't gzip a jpeg since it's already compressed.
All Accept-* headers are like this. e.g. the server doesn't necessarily support any of the languages requested in the Accept-Language header, but it doesn't hurt to ask. You always have to inspect the response headers to see the result of negotiation.
Similar thing is with header that's quite useful, but for some reason very few sites honor it: "Accept-Language" browser can specify which languages are preferred, but it is up to server to honor it (for example given language version is not available).
> The file size must be between 1,000 and 10,000,000 bytes.
[1]: https://docs.aws.amazon.com/AmazonCloudFront/latest/Develope...
The docs do not explain why.
As everything will end up in a packet when sent through the network stack, you might want to choose your minimum input size in such way, that you generate gzip-compressed output big enough. Why big enought? Nagle's algorithm [1]
So yet another reason to think about 'what to gzip'.
More sophisticated algorithms can either decide exactly which packets to send, or use TCP_CORK to shove part of a packet into a buffer before they add the rest of the stuff, e.g. preparing HTTP headers and then adding the static document that goes after them.