Imagine a data stream X that is compressed into another data stream Y. Imagine that a small portion of X is data stream x1 which is the portion of data used for a signature. That will get compressed into y1. Now lets define x2 as all the data in X that is not x1. Now if you are always guaranteed that the same x1 would get compressed into the same y1, then things would be easily predictable and you can just compare compressed signatures. But this is not the case. If x2 is different, then the same x1 can be compressed into a different string.
What's probably happening is virus scanners are taking the easy way out rather than implementing decompression algorithms in their scanning engines (it probably looks better on traditional benchmarks too)
On your second point I agree. I'll refrain from ranting about AV programs here, but suffice to say there hasn't been enough innovation in the field because AV companies are able to sell substandard products and still make good money.
The gzip format documentation is available here btw: http://www.gzip.org/zlib/rfc-gzip.html