Intel QuickAssist Technology Zstandard Plugin for Zstandard
community.intel.com
community.intel.com
That's not to detract from everything AMD has done, but hardware is only the first step. Software that properly uses the features your hardware provides is just, if not more, important.
I love the fact AMD is pushing Intel so much. Pre-C2D days were amazing because we had two vibrant, innovative companies pushing to the edge of possible; trying to out-do each other. Pre-Ryzen was a horrible time. Do you want to spend $500 to upgrade from a 4-core intel 4000-cpu to an intel 5000-cpu? You'll get DDR4 and 1% IPC.
Now we get massive IPC, clock speed, ram and PCIE improvements on a regular basis. Competition is great, especially for the consumer.
Sure, benchmarks of pure compression will come out ahead for Intel, but a typical sever running something like a database engine with compressed disk storage is likely a win for AMD.
The new EPYC 9004 chips are much cheaper than the Intel CPUs with QAT, and they have more cores, and much higher clock speeds.
It’s hard to believe there’s any scenario where the Xeon CPUs provide better overall value…
https://bugs.chromium.org/p/chromium/issues/detail?id=124697...
Relatively few servers run databases. Many more are running web servers and given the huge efforts web devs go to in order to optimize response sizes, much faster, lower power and lower latency compression seems like an instant win. All that's required is for web servers to integrate the zstd library and this QAT thing, and that can be enough to tip the balance especially for IO bound servers that are mostly just doing string interpolation and waiting for backends.
And you mention a DB with compressed disk storage. How is better compression not a huge win for that use case? Databases are usually disk IOP, latency and CPU power constrained, and this is a win for all of those cases.
Finally, clearly this tech can be applied to other algorithms not just zstd. Presumably they highlight zstd because the library is actively maintained and was willing to add this (probably single use) "plugin" API. Really they should have just integrated it directly instead of complicating things for every zstd user, but I guess it adds extra dependencies or increases code size or something. The wins are big enough that other codec libraries will probably use the same approach and eventually it'll be abstracted by APIs that are less code size sensitive.
Yes, AMD is doing great right now, but Intel have had an edge when it comes to specialized CPU features for a long time. For example AMD are only just now catching up to where Intel were with SGX 8 years ago.
Not parent, but I guess what he / she meant was total workload. E.g, If 100% of your workload, only 5% of that is compression, speeding it up by 5x will only give you a 4% reduction in total system workload.
Which means other factor like clock speed, more core, cache etc in a different system could still win despite it lost the compression benchmarks.
Ah yes, being hopelessly broken?
But that was just an example. Intel have introduced a quite a few interesting CPU extensions and special bits of hardware that AMD had no answer for over the years.
- CDNs, webservers and/or load balancers: SSL termination and compression acceleration for serving content
- VPN nodes: Encryption acceleration
- Storage: Compression acceleration
It's sad they stopped making PCIe cards, and the older PCIe cards aren't compatible with the new drivers and "hardware revision". You could combine AMD processors with Intel QAT acceleration.The major downside of course is it is quite tricky to use this stuff in practice. In the cloud, you need a bare metal instance that exposes the QAT peripheral, and they are relatively scarce. And this whole generation of hardware is only just beginning to land in public clouds. For machines you own, you will need to scrutinize Intel's somewhat ridiculous product matrix in order to acquire a Xeon that has QAT.
Is gzip actually obsolete or are there just newer alternatives? gzip is still everywhere
That said, the igzip re-implementation is really good. If you can use it, igzip moves gzip closer to the optimal frontier.
Plus, if you want super compatibility there is older-than-dirt pkzip/infozip.
But yes, lz4 and zstd would be what I'd recommend people use.
zlib-ng tried to merge as much as they could from the cloudflare fork without getting into licensing issues.
I would just go for lz4 and zstd, really, at this point, but if for some reason gzip is necessary or fits your use case better, ...
(Though I'd really like to hear what that is.)
Zstd can use a much larger window (8MB recommended) and a much better entropy coder: https://github.com/Cyan4973/FiniteStateEntropy
Now you know why they are comparing 16KB payloads only.
We also have some half baked ideas to combine a fast SW match finder that only looks for matches >64KB away, and supplements the matches that QAT finds.
Download TurbBench from Releases [2]
Here Some Benchmarks:
- https://github.com/zlib-ng/zlib-ng/issues/1486
- https://github.com/powturbo/TurboBench/issues/43
So I think good acceleration for things like compression is going to be a big help.
https://github.com/intel/QAT-ZSTD-Plugin/blob/main/src/qatse...
I don’t fault Intel for choosing web frontend acceleration over storage first, but this has been a long time coming.
Dealing with the gzip core vendor and the FPGA vendor (both in wildly different timezones) was a little unpleasant.
For example, the Zstandard plugin has this requirements:
Hardware Requirements
Intel® 4xxx (Intel® QuickAssist Technology Gen 4)
Software Requirements
ZSTD* library of version 1.5.4+
Intel® QAT Driver for Linux* Hardware v2.0
What's Gen 4? (I later found that means '4th Generation Intel® Xeon® Scalable Processors'. But not any Gen4, you have to find one with QAT included, and the product matrix is huge.)What's the v2.0 of the hardware? (I'm not completely sure about that)
Does it automatically accelerate zlib and openssllib calls, or do you have to patch them? (I think you have to patch them)
I think they're stabilizing the API from now on, but it's not as simple as buying a CPU or one of the older PCIe cards, and loading the driver.
[1] https://github.com/intel/isa-l/tree/master/igzip [2] https://github.com/mxmlnkn/rapidgzip [3] https://github.com/intel/QATzip#test-qatzip
> For the Silesia corpus, data compression ratios are:
QAT-ZSTD level 9: 2.76
zstd level 4: 2.74
zstd level 5: 2.77
This presumably also means the best possible compression that QAT can achieve is worse than what vanilla zstd can do.QAT cannot compress across blocks. So as soon as your input is more than 128KB the compression ratio tanks.
Usually well below what even level 1 of software zstd does. They choose the input very carefully.
edit: CPA_ACC_SVC_TYPE_XML https://github.com/ravynsoft/ravynos/blob/ee81203faa28ab3e6a...
It also seems to suggest that it could accelerate hyperscan, but considering the hardware and software are both from Intel and they don't integrate them, maybe it doesn't work or is theoretical.
In that context, XML (presumably parsing) is not a surprise.
Best way to optimise the reusable software is to turn it into single hardware CPU instruction
At "only" 100GBPs per adapter, with said adapter costing like 1 Epyc, does it make sense?
Epycs can do 400gbps of compression in software, without much SSE, and handwritten assembler.