Sure - it's a handy tool. But at scale surely someone had enough time to research better options? And at Amazon scale, it probably even pays to hire a team to write a custom compression algorithm perfectly tuned to your compute and storage?
Sure - it's a handy tool. But at scale surely someone had enough time to research better options? And at Amazon scale, it probably even pays to hire a team to write a custom compression algorithm perfectly tuned to your compute and storage?
More importantly, it would require researching whether something even better might be just around the corner. I don’t think Amazon wants to switch compression every few months.
In summary, lower the bar for trying new things. This is how we innovate.
I found out the way to do it was to have Nginx run a proxy server that caches the output of another Nginx 'server' that does the Brotli dialled up to 11. So you are making your own CDN.
Nothing is difficult when you know it, but Gzip has scratched the compression itch so well that people just do not have a problem with it and therefore do not seek to change it.
It is almost like encryption in this regard.
When you look at it like this, not very surprising that initiatives like cost savings optimizations may take a back seat for periods of time.
The short answer to "Why still using gzip in 2022" is almost always answered by "Because it made no business or financial sense to spend head count on it".
Amazon is usually fairly smart about the ways it spends head count, particularly on things that could notably reduce costs. Managers/directors/VPs/SVPs obsess over what the value is from various work, and routinely adjust priorities and reallocate head count. They track not just what is happening within individual services, but also the overall business strategies, what's on the horizon etc.
One simple example that most folks outside the industry won't be seeing is that every major cloud has been working their collective arses off for well over a year on meeting JWCC contract needs. There's a lot of work involved in that contract, because unsurprisingly it's really hard to build and run an entirely air-gapped cloud region, but the payoff on being selected as a vendor is phenomenal. Far more than saving an additional 30% of storage, even at AWS S3 scale. _It's not the only major business and engineering initiative_ that will be taking place across AWS.
I've got a long list of "We should do x, y, z" that applies to my service in the cloud I work for, a number of which will see notable performance improvements. They make zero sense allocating head count to, though, because I've got at least a dozen other more important things going on from a business roadmap perspective to get solved that will make us way more money and/or save way more engineering resources down the line, be it automation stuff, or new features.
What does happen, though, is that list is kept in mind any time new business priorities come about. If there is any way bits on the list can be tied in to something the business wants, it'll get hooked in to it. The business gets what it wants, the service gets what I want, everyone is happy.
(To repeat something others have pointed out, there's also zero indication of timeline, it could have happened any time in the last several years that zstd has been a thing, they may well have switched to it almost as soon as Facebook removed that ridiculous "You can't sue us" clause)
Let's say you save a few hundred million dollars a year after switching. The cost to switch is almost certainly more than a few hundred million dollars. When you make that kind of investment—especially when you're moving away from a very boring technology—you want to be damn sure you know exactly what you're getting yourself into.
Second I believe youre underestimating the scope of the problem. Amazon has thousands of teams, even more services, each with their own priorities, and innumerable different access patterns, data types, etc. For all practical purposes there is no single “they.”
To get an idea of the scope how long would it take you to remove gz from oh … 50,000 projects? Each with many deployments and up to 15 years of active data.
https://pure.tudelft.nl/ws/portalfiles/portal/94907930/Jiany...
If it is far enough above the curve of speed Vs ratio, then it will pay for itself in saved storage.
Remember there is no need for said FPGA to be in every machine - in a data center with ten's of gigabits of bandwidth to every node, you can send data to another machine for compression and receive back the compressed data to store.