AWS switch from gzip to zstd – about 30% reduction in compressed S3 storage
twitter.com
twitter.com
I don't think AWS routinely compress customer data that I've noticed. I guess he must mean for their internal products that use S3 perhaps?
It might make more sense on their "deep archive" product maybe where the customer has to commit to a minimum storage retention and also pay a retrieval charge which scales with the amount of data recovered (hence paying for the CPU to decompress).
If you can stuff more bits in an IO operation, you’re winning.
Modify the compression level to try and keep the output buffer at 60% full.
On Linux PSI ( /proc/pressure/io ) probably provides more accurate information (the code already uses /proc/cpuinfo ).
Detail: in the fileio.c module there are lines such as:
if (oldIPos == inBuff.pos) inputBlocked++; /* input buffer is full and can't take any more : input speed is faster than consumption rate /
if ( (inputBlocked > inputPresented / 8) / input is waiting often, because input buffers is full : compression or output too slow */
This impacts a 'speedChange' variable.
Its potential values (an enum) are 'noChange', 'slower', and 'faster'. They are processed rather simply: if (speedChange == slower) { ((...)) compressionLevel ++;
if (speedChange == faster) { ((...)) compressionLevel --;
For any storage system like this you usually have a few bottlenecks. IO and Network are the obvious ones, followed by tiering (cache, fast io, slow io, ...) and at the very end CPU.
Now let's say network is your bottleneck. If you can send the data to the client in a compressed for then you get the compression ratio as additional bandwidth. And the user would get the data quicker! So compression to the network is a clear win.
But the common bottleneck is often IO, a high end SSDs with 1M IOPS at 4KB would _theoretically_ serve 4GB/s, a 40GBit link. That's without any redundancy over other overhead.
Again compression to the storage layer would decrease the total amount of IOs, thus making sure a customer gets data quicker.
Ok, let's say both are not the issue. The fastest compression algorithms compete with memcopy. So if you need just one copy of your data you might have been faster by compressing it.
Especially fast compression algorithms (zstd, lz4, snappy, lzo, ...) are worth the CPU cost with virtually no downsides. The problem is finding the right sweet spot that reduces the current bottleneck without creating a CPU bottleneck, but zstd offers the greatest flexibility there, too.
Oh for range requests.... Those large objects are likely split anyway, for easier error recovery (imagine 100MB into a 1GB transfer you notice that the file data was corrupted - not good). Once you work on blocks it's easy to do somewhat efficient range requests again.
That said, a surprising number of video sources use sparse file type setups, and I've gotten pretty good compression (up to 60%) using LZ4 with NVR files from some brands.
But we are talking about multi-tenant systems here. You will get better performance and/or better prices if the system finds compressible data for other customers.
All the benefits hold, even for you with non compressible data, if there is an overall benefit.
Hitting a bottleneck less often then this will reduce your tail latencies, too.
It is just that _your files_ won't contribute to these improvements. But you get all the benefits as well.
It would be waste to _not_ use the CPU time.
Granted, the CPU powet consumption would drop with enabled power management, but that's only marginal gains as the CPU is likely busy anyway and the total duration of busy time might drop (race to sleep)
I have used this inverted logic to great extends. E.g. when expanding data center capacity I was usually swapping old harware for new hardware once it arrived. It is way better to have the older generations as spares and use the new capacity at the most utilized places.
Do not waste the things you pay for!
Hardware compression is sadly not widely available, I think the only consumer product I know with it is the PlayStation 5.
The mainframes from IBM have had hardware zlib since Z14 iirc and in my small tests it is very fast compared to the CPU implementation
I would expect this to become part of newer generations of CPUs once it becomes popular.
My Intel i9 11900k seems to support it or has it built-in.
Example:
You have an existing 40 gigabyte file
It happened to compress well
You delete it and your free disk space goes up by 4 gigabytes.
You then write a new 40 gigabyte file that doesn’t compress well
Replacing an existing file of the same size just ate an extra 36 gigabytes.
How would you plan around that? SSDs should store the bytes given and don’t play fancy games.
Given that, it's good sense to compress if at all possible simply to make the drive live longer.
And guess what - I just googled, tons of hits, and this has been done for a long time :)
So it makes sense, is done, and is important for modern SSD behavior.
I'm surprised (and shocked) that letting unencrypted data hit the disk is still common enough to make such optimizations worth it.
Even if you just stick the key in the server's TPM without any sealing, an encrypted disk makes it much easier to deal with e.g. drive returns (for warranty or fault analysis) or disposal.
There's very, very little benefit in encrypting data at a filesystem level in a datacenter if you think about it.
"Even if you just stick the key in the server's TPM without any sealing, an encrypted disk makes it much easier to deal with e.g. drive returns (for warranty or fault analysis) or disposal."
Would you resell usable drives that you no longer want to use (e.g. because they're too small or your needs changed more towards SSDs) if unencrypted data was written to them at some point?
How much effort would you put into making sure that no broken drive that can't be wiped leaves the datacenter without shredding?
Any process you put in place will have gaps - e.g. through human error or malicious acts - and this provides pretty solid protection against that.
And unfortunately compression and encryption are seemingly at odds fundamentally :c
Bitlocker supports on drive hardware encryption, and I'd be surprised if other major file systems didn't.
If I recall, it's a FIPS requirement for data at rest now.
Because the storage hardware might be loyal to someone who is not the owner of the data.
Because the hardware cannot be trusted to do it correctly. IIRC Bitlocker stopped relying on it for this reason.
https://www.howtogeek.com/fyi/you-cant-trust-bitlocker-to-en... (see the updates)
>Data compression via encoding algorithms enables a solid state drive (SSD) to write less data, which in turn yields higher write bandwidth. With a significant amount of data being compressible, performance benefits can be substantial.
https://www.intel.com/content/www/us/en/support/articles/000...
Adding life to SSDs is a terribly useful feature
Maybe they charge at uncompressed rate but store it at compressed? Then they got even more money!
Only temporarily with SSDs. With spinning rust, it also often paid off to compress data. We'd store large treebanks compressed, because decompression was much faster than disk reads.
Since CPUs are fast enough to deflate in real-time now, your bottleneck for a read is your storage/network.
Reducing the bytes read from storage improves the IO latency.
this is pretty easy, you flush the compression buffer every megabyte or so and maintain an index
maybe 50 lines of code
I implemented this as part of the MinIO server. See "Seeking Compressed Files" here: https://blog.min.io/transparent-data-compression/
We choose a compressor without literal compression for a faster baseline, but the concept remains the same.
the index goes in there, no extra seek needed
yeah, 4 bytes for every megabyte
> s3 really isn't the right layer to implement compression. filesystems aren't either. it's better to leave it up to the application.
yeah, I'm sure you're right and Amazon have absolutely no idea what they're doing and like to spend unnecessary CPU cycles doing pointless work and add "significant bloat" to their metadata
... or, you're wrong (like in every previous comment in this chain)
this tweet is not talking about compressing customer data in s3, i seriously doubt that aws compresses customer data in s3 for all the reasons i've already listed. i am right and amazon does know what they're doing, which is why they don't compress customer data in s3.
4 bytes per megabyte becomes significant at scale when you have to keep it in ram, which you have to do if you want to avoid the extra IO.
Each part can be max 5GiB as per S3 spec. 5120 * 4 = 20KiB.
Even if you unpack to 8*2 bytes in memory when decoding, you are still not talking a huge amount of memory.
The on-disk space is ~0.0004% as blibble calculated, and should easily be offset by the compression achieved. In MinIO we don't store indexes for files < 8MiB, so for small files there is no overhead.
If the added metadata is a problem for whatever system you are looking at, then that is a characteristic of that system and not a general problem.
and you don't understand the algorithm if you think you need to keep the index in RAM, because you don't
as has been explained to you several times, you don't
> if you pollute the metadata cache with useless junk like this
the overhead is 0.0004% with 1 index entry per megabyte, and if that's too much that can be reduced by 10/100/1000/10000x that by changing the size
as we're clearly now going around in circles, I won't be responding again.
That's not a good argument: they could lower their costs with compression, still charge the same, and make more profit.
If the difference between standard S3 and S3 Glacier was just slower disk, then rate limiting the customer would suffice.
But if there’s a significant amount of compute thrown at data de-duplication, compression, and indexing, then it starts to clarify why there’s a pricing penalty for using Glacier with the same access patterns as one would use on standard storage.
And what if they can charge you the uncompressed size and only actually store much smaller compressed files behind the scene. That seems like having your cake and eat it too.
Like when it was easy to file share on dropbox by having the correct hashes. A GUID could summon a 1GB file.
Content addressing across accounts on private, AWS encrypted S3 buckets would run counter to their claims.
tl;dr:There's a real efficiency/security/insider risk trade-off here.
Edit: I should disclose that I work for a competitor. Don't intend any astroturfing.
That said, of course, customers can upload encrypted blobs of uncompressed data. But I’d call it an exception that proves the rule. Here service simplicity should win and those blobs may end up recompressed.
Since we already used a Snappy-derived method, each 1MB block is stored without backreferences. With this we only have to decode at most 1MB-1 extra bytes to respond with a specific range offset.
Maybe data in S3 is compressed? I don't know the internals of S3.
You may have switched transparently to using it (e.g. for your application log files), via internal Amazon tooling (that many AWS services use).
However, this tweet wasn't about AWS customer's data, but AWS's own data.
> Maybe data in S3 is compressed? I don't know the internals of S3.
Almost every other AWS service uses S3 for storing something.
Think of EBS snapshots, DynamoDB backups, RDS backups.
And that's not even the service-internal data.
AWS switched our own service's log storage (in S3) from gzip to ztsd a bunch of years ago -and reduced storage costs by 30%.
This is AWS's service data (which is still exabytes) and AWS realized the efficiency.
So all the comments of "where is my 30% discount" are offbase.
Source: I've worked on AWS for a bunch of years.
>FSE, short for Finite State Entropy, is an entropy codec based on ANS. FSE encoding/decoding involves a state that is carried over between symbols, so decoding must be done in the opposite direction as encoding. Therefore, all FSE bitstreams are read from end to beginning. Note that the order of the bits in the stream is not reversed, we just read the elements in the reverse order they are written.
A) Use arithmetic coding to encode symbols as fractions instead of using an integer number of bits for each symbol. https://en.wikipedia.org/wiki/Arithmetic_coding http://www.ws.binghamton.edu/fowler/fowler%20personal%20page...
B) Shuffle the bits around so that you can use faster math and lookup tables instead of directly operating on fractions.
He's talking about runtime logs of the S3 service, or possibly logs from Amazon's services that are written to S3.
Amazon has ~100K services running on millions of EC2 instances. (Generalizing) each service writes logs to the local machine, and then later compresses and transmits them to a set of systems that are backed by S3 for monitoring / further bulk processing etc.
Some back of the envelope math on this:
For 1 exabyte to be the number, that equates to < 1TB average logs kept per host. Depending on what the service is doing log retention is 3mo / 1y / 10y. Using those numbers we're talking per service average 12-460 MB/hour.
https://old.reddit.com/r/programming/comments/wtd61q/aws_swi...
Sure - it's a handy tool. But at scale surely someone had enough time to research better options? And at Amazon scale, it probably even pays to hire a team to write a custom compression algorithm perfectly tuned to your compute and storage?
More importantly, it would require researching whether something even better might be just around the corner. I don’t think Amazon wants to switch compression every few months.
In summary, lower the bar for trying new things. This is how we innovate.
I found out the way to do it was to have Nginx run a proxy server that caches the output of another Nginx 'server' that does the Brotli dialled up to 11. So you are making your own CDN.
Nothing is difficult when you know it, but Gzip has scratched the compression itch so well that people just do not have a problem with it and therefore do not seek to change it.
It is almost like encryption in this regard.
When you look at it like this, not very surprising that initiatives like cost savings optimizations may take a back seat for periods of time.
The short answer to "Why still using gzip in 2022" is almost always answered by "Because it made no business or financial sense to spend head count on it".
Amazon is usually fairly smart about the ways it spends head count, particularly on things that could notably reduce costs. Managers/directors/VPs/SVPs obsess over what the value is from various work, and routinely adjust priorities and reallocate head count. They track not just what is happening within individual services, but also the overall business strategies, what's on the horizon etc.
One simple example that most folks outside the industry won't be seeing is that every major cloud has been working their collective arses off for well over a year on meeting JWCC contract needs. There's a lot of work involved in that contract, because unsurprisingly it's really hard to build and run an entirely air-gapped cloud region, but the payoff on being selected as a vendor is phenomenal. Far more than saving an additional 30% of storage, even at AWS S3 scale. _It's not the only major business and engineering initiative_ that will be taking place across AWS.
I've got a long list of "We should do x, y, z" that applies to my service in the cloud I work for, a number of which will see notable performance improvements. They make zero sense allocating head count to, though, because I've got at least a dozen other more important things going on from a business roadmap perspective to get solved that will make us way more money and/or save way more engineering resources down the line, be it automation stuff, or new features.
What does happen, though, is that list is kept in mind any time new business priorities come about. If there is any way bits on the list can be tied in to something the business wants, it'll get hooked in to it. The business gets what it wants, the service gets what I want, everyone is happy.
(To repeat something others have pointed out, there's also zero indication of timeline, it could have happened any time in the last several years that zstd has been a thing, they may well have switched to it almost as soon as Facebook removed that ridiculous "You can't sue us" clause)
Let's say you save a few hundred million dollars a year after switching. The cost to switch is almost certainly more than a few hundred million dollars. When you make that kind of investment—especially when you're moving away from a very boring technology—you want to be damn sure you know exactly what you're getting yourself into.
Second I believe youre underestimating the scope of the problem. Amazon has thousands of teams, even more services, each with their own priorities, and innumerable different access patterns, data types, etc. For all practical purposes there is no single “they.”
To get an idea of the scope how long would it take you to remove gz from oh … 50,000 projects? Each with many deployments and up to 15 years of active data.
https://pure.tudelft.nl/ws/portalfiles/portal/94907930/Jiany...
If it is far enough above the curve of speed Vs ratio, then it will pay for itself in saved storage.
Remember there is no need for said FPGA to be in every machine - in a data center with ten's of gigabits of bandwidth to every node, you can send data to another machine for compression and receive back the compressed data to store.
That being said, looking at things from the outside, I tend to think that objects are not stored in a compressed form. It would make many operations much slower.
1. There are APIs to fetch byte ranges from objects. Would that work without decompressing the entire object?
https://docs.aws.amazon.com/whitepapers/latest/s3-optimizing...
2. You can query over multiple objects using SQL either via Athena or S3 Select.
https://docs.aws.amazon.com/AmazonS3/latest/userguide/s3-gla...
Edit:
Someone has already explained how byte range requests would work.
I wonder if you could store the objects in chunks and then individually compressed. Create an index of the byte ranges for those chunked objects so they can be looked up easily.
For example:
range 0-4000 -> objectA, objectB, objectC
Then just return those 3 objects, decompressed.
Where?
This isn't about them shrinking your individual files, its compression on their internal transport.
I suppose you could speculatively decompress and then re-compress and see if you get the original compressed file back, and maybe most people happen to use the same compression implementations with default settings.
EPUB is an example of this. They're mostly bog-standard ZIP archives, but in order for the file to be valid, the "first" file in the archive, linearly speaking, must be named "mimetype" and stored with no compression (each file in a ZIP archive can have a different compression level). If you just unzip an EPUB and then just dumbly zip it back up again, your end file will not be a valid EPUB.
This puts a serious limitation on your compression. You would only be able to re-do the entropy coding part of DEFLATE, which is actually pretty good.
You would still need to store the original Huffman tables for each block, so you can reconstruct the entropy coding exactly.
I doubt this would even gain you a single percentage.
[1] https://github.com/deus-libri/preflate
[2] https://github.com/schnaader/precomp-cpp/ (which internally makes use of preflate)
I doubt Amazon intends to bill on compressed size.
I get hashing it and providing hashes, but encrypting public files seems excessive.
I've seen a pattern of: drop raw data into an S3 bucket that has a very restrictive policy with a long retention policy. Then, process that data asynchronously (encrypt, transform, filter, etc) and drop it into a different bucket/area that is accessed by other consumers.
Then, if any part of your ETL fails (encryption included), you can fix your bug and reprocess from your raw data without writers seeing any impact.
Did some searches but came up with some guesses like https://maisonbisson.com/post/how-big-is-s3/ but they are dated and still just guesses.
On pure storage that is 100,000,000TB. That is 1 billion dollar already excluding any redundancy required by S3.
It’s really remarkable for my use case. I compress gigabyte-sized SQLite databases under the GitHub 100 MB file limit since Git LFS is a non-starter. My tables are collections of follower and stargazer counts to measure user and repository popularity and understand where one might be on a distribution compared to a subset of other GitHub data points.
[1]: https://github.com/andrewmcwattersandco/github-statistics
gzip will stay with us because of ubiquity, just as we mostly use image formats from the dot-com era. But if you control both sides, zstd is a big upgrade.
gzip compression levels are almost useless and typically result in very little actual compression ratio differences, but typically with massive CPU usage differences.
zstd doesn't do everything better than gzip though; I believe gzip still offers somewhat faster compression for equal compression ratios last time I looked (but decompression is much faster). I mostly replaced gzip with zstd myself, but there are still scenarios where gzip might be preferable.
Compression time CPU is lower for equivalent ratios. Maximum compression is better. Decompression CPU is much lower (i.e. faster), and that is independent of the level used at compression time.
Overall, zstd still came out as the clear winner (also compared some other compression tools), so I went with that.
It would be a reduction in data, but not compressed data.
https://github.com/minio/minio/blob/master/internal/s3select...
For RAW --> lossy format, storing them in JPEG-XL would be preferable, and I doubt most compression algos would do better than JPEG-XL.
Depends on the data - you can have totally different output compressed archive if compare to winrar for example.
Example: I doing backups time to time, and archive important files into compressed archive containers (rar / 7zstd).
I've noticed, that my dev folder with tons of repositories, images, and different work related fines - vary damn too much.
/dev/ size = ~11GB
rar output (normal compression) = ~2.1 GB zstd archive output (normal compression) = ~4.7GB
Why? linked files, same files not treated as a 1 file + links to these files. Instead these files compressed each 1 by 1 instead of copy 1 identical, and compress the file. And many things like that.
Suggestion for 7z-zstd -> add ability to save links to files, not treat them as separate files, and adding an option like in winrar to search for identical files first and re-link all of them and remove duplicates, instead of compression each.
Archival software then has to build a file format on top of the compression algorithm, and there are multiple ways to slice the problem. For example, a tar.gz will first tar everything into a big archive file, then feed it into gzip for compression. zip, on the other hand, feeds each file individually into the chosen algorithm (DEFLATE for most implementations).
Your critique is that the 7zip archive format is not suitable for use with zstd in the case of many small yet identical files. zstd is doing its job, just the archival format is not playing along.
Of course there are cheaper cloud services like B2 at $5/TB/mo, but even that can’t compete on price with physical infra when you have more than a handful of drives.
$15/TB is hard enough for a home buyer to reach for hard drives, excluding server cost. Where did you get $10?
$9.70/TB 10TB: https://www.ebay.com/itm/275400447467
$8.31/TB 4x8TB: https://www.ebay.com/itm/125132232253
You can get even cheaper if you buy in bulk.
Server cost is $1/TB with reasonable density
It’s econ101, the cost of providing the service doesn’t really factor into the price charged, what does factor in is the cost of alternatives (storage from azure, backblaze, on prem, not doing it), and the risks (increase s3 costs and people might realise they are being ripped off on ec2)