How to get your backup to half of its size – ZSTD support
percona.com
percona.com
Results shows that ZSTD not only overcame LZ4 results on all tests, but it also brought backup size to half of its original size.
When streaming is added to the mix is when we see the biggest difference between both algorithms, with ZSTD overcoming LZ4 with an even bigger margin.
This can bring users and organizations a huge amount of savings in backup storage, either on-premises or especially in the cloud – where we are charged for each GB of storage we use.
The interesting thing in comparing compression algorithms is — "what are you optimizing for?" Compression speed? Decompression speed? Storage savings? There are Tradeoffs. Getting the most storage savings might be suboptimal with a database if it adds too much latency or hampers throughput.
Nicely done, Percona!
https://www.scylladb.com/2019/10/07/compression-in-scylla-pa...
zstd also can tune itself and adjust ratios to saturate output bandwidth, which is pretty cool (--adapt).
These XML files are full backups rather than deltas - so each one contains the full previous file plus the new additional data.
My assumption was if I grabbed just a swath of the old ones, say a years worth and compressed them together I would get really decent savings.
I was pretty disappointed - gzip knocked 10 gb off of about 100 gb of data. I started doing some research and found people saying 7zip and it's sliding dictionary size options were the answer. After multiple tries, each run of 7zip taking multiple days I was able to get it down to about 70 gigabytes from 100. Better then gzip but frankly nowhere near what I would expect.
Does there exist a compression that could better handle this sort of expanding documents?
Is the XML wrapping a bunch of other random data or something?
As an example, I have zstd enabled on some zfs pool. The client-side encrypted time machine backups does not even compress 1 %, as expected.
You can’t embed binary data in xml raw, so a common pattern is embedding it as base64, or similar type of wrapping.
Technically, it’s possible to escape it in other ways, but it’s error prone.
Either way, XML which has text strings or whatever typical XML document data that ISN’T something like base64 encoded random data should compress dramatically better than what the poster was talking about.
So, what the hell is in your XML anyway?
Some compressors have options for solid compression, or you can use tar to first concatenate the files. For both, you need to sort the files first, so that they are compressed in chronological order, otherwise the common information will drop out of the dictionary.
And how big is one file and how much RAM you have? Because you might need increasing the "dictionary size" of the compression algorithm, both 7zip and zstd support multi-gigabyte sizes.
I've got 64gb of ram and 10 cores to work with, but it seems single core compression might be in order?
Yes: storing deltas. Git might work well. A backup system that does subfile deduplication may work too (e.g. Restic, Borg).
I was under the impression that git didn't do well with large files.
I've used it to store 20-80GB of files in a repo and it works quite well.
Long Range ZIP or LZMA RZIP
https://github.com/ckolivas/lrzip
"A compression utility that excels at compressing large files (usually > 10-50 MB). Larger files and/or more free RAM means that the utility will be able to more effectively compress your files (ie: faster / smaller size), especially if the filesize(s) exceed 100 MB. You can either choose to optimise for speed (fast compression / decompression) or size, but not both."
Impressive!
This is like the claims of tape drive manufacturers who claimed double the capacity of the media.
Pure marketing wank.
If you're not already compressing and encrypting your backups, then you're not doing it right to begin with. zstd may not be the most efficient algorithm for a given dataset. There are hundreds of compressors out there. It's worth finding the best one on average for a given job and occasionally revalidating it against other candidates.
PSA: Untested backups aren't backups.
Maybe React too, I'm not a front-end person.
Cons: created a post-truth world which destroyed democracy and killed millions of people by spreading antivaxx propaganda
Eh.
I can't think of any real pros to Twitter. Did they create any great new open-source technologies?
Now Airbnb, those people caused untold suffering by both their impact on housing market, and by creating Airflow ...
You might as well be blaming highways and radio - Facebook made it more accessible, but far from created it.
The game isn't over yet, however.
We actually went back (SQL backups, not xtrabackup) because weirdly enough bzip2 made smaller files when we aimed for compression ratio with similar compress time as bzip2. THink zstd got a bit faster since then tho...
Only for MySQL, PostgreSQL were noticeably smaller with zstd