To put that in perspective, consider that a large portion of the file is literally a nucleotide sequence, e.g. "GGGTGATGGCCGCTGCCGATGGCGTCAAATCCCACC" and that the rest of the file format is also fairly repetitive. If we just imagine only the nucleotide sequence, one way to compress that would be to take advantage of going from an 8 bit alphabet (ASCII encoding) to a 4 bit alphabet (CGAT). Naively, you'd expect ~50% compression from that alone, and be left with a bit stream that's still compressible.