Reverse Engineering the Comtech AHA363 PCIe Gzip Accelerator Board
tomverbeure.github.io
tomverbeure.github.io
https://www.businesswire.com/news/home/20081008005207/en/Com...
http://www.aha.com/DrawProducts.aspx?Action=GetProductDetail...
That still exists for Allwinner CPUs. You can offload the TLS handshake.
https://blog.cloudflare.com/on-the-dangers-of-intels-frequen...
first off, AVX2 is a subset of SIMD instructions - SIMDjson includes a large number of implementations including several which do not use AVX2 at all.
next, only a subset of AVX2 instructions cause downclocking (namely, the 512 bit wide ones), which SIMDjson does not use. throwing out the use of "SIMD" entirely because you read an article about a specific AVX512 instruction which caused down-clocking might be a bit premature, no?
https://www.snia.org/sites/default/files/SDC/2019/presentati...
> The input blocks, while compressed independently, have the last 32K of the previous block loaded as a preset dictionary to preserve the compression effectiveness of deflating in a single thread
(Note that the heatsink is on the FPGA, not the actual gzip accelerator chip.)
The normal development flow for FPGA software is similar to ASIC in that people focus on testing as many elements as possible in software simulation to a very high level of coverage before even downloading to the FPGA. Once in there, you're reliant on JTAG (high-speed serial bus) to read out values from the target device.
Tools like ChipScope can let you see what's going on and set ""breakpoints"". http://web.mit.edu/6.111/www/labkit/chipscope.shtml
All of this is much harder when it's on a board you didn't design!
One thing that didn’t make it in the blog post was that a strategically soldered wire shortened the JTAG chain to only include the 2 Intel chips and bypass the AHA chips.
The power to the AHA chips gets cut off at some point, breaking the JTAG chain, but the FPGAs stay active.
So by bypassing, you keep the chain alive.
It might be pretty nifty to run swap through this...
PS5 has a IO chip and extra architecture for decompressing textures (here the context with this gzip accelerator) and more features.
Mark Cerny said that they needed this chip because from a CPU ressource usage,it would use all CPU cores. So nvm now are so fast that a io co processor is feasable again.
> IO chip and extra architecture for decompressing textures
you dont want uncompressed textures in your GPU memory, compressing DXT compressed textures is not idea to say the least, you can count on 30-50% compression ratio, not the marketing 2x peak number Sony was throwing around
>would use all CPU cores
is marketing exaggeration, LZ4 decompression can achieve ~3GB/s per core. How it is today: https://www.jonolick.com/home/oodle-and-ue4-loading-time
My feeling is, that what Sony is now doing, goes in that direction.
I'm highly curious and hopefull.
I will read your posted blog, looks interesting. Lets see what will arrive in the industry at the end of the day.
Also, AHA has a followup product that doubles the performance by dropping 4 ASICs and 1 FPGA on a board instead of 2 ASICs. So modularity is a factor as well.
But I think the first point is very likely the reason.
the FPGA does the PCIe and all necessary data marshaling, which allows a lot of flexibility for updates and bug fixing on the most finicky parts of hardware
two ASICs because once you have one manufactured, most of the costs are sunk and per unit it's cheap to stick a second on the board