100Gbps RF Sample Offload for RFSoC Using GNU Radio and PYNQ
strathprints.strath.ac.uk
strathprints.strath.ac.uk
That’s absolutely nuts. The board is a bit over $2k, and they don’t go into how effective the RF front end on that board is, but enabling that kind of bandwidth at that price is amazing.
> The RFSOC 4x2 is currently only available to Universities and Research Institutes.
And you're almost right. RFSoC XCZU48DR-1FFVG1517E costs $22742 (not in stock) and XCZU48DR-2FFVG1517E costs $31840 (not in stock).
* DisplayPort 2.1
* USB4 version 2
* Thunderbolt 5
Amazing for desktops and notebooks.
Hopefully such boards could also make soon into at least SDR hobby price range.
Is this useful for anything else than visualization of the spectrum in a waterfall diagram?
[1] Ultra-wideband SDR architecture for AMD RFSoCs using PYNQ based GNU Radio blocks:
GPUs are incredibly versatile. As CPUs struggle to add more cores instead we should shift software as much as possible on to the GPU and use the CPU as a coprocessor.
At least the GPU should be given direct access to storage.
Perhaps GPUs should come with 1-2 NVMe slots and upgradeable VRAM.
also, LOL at nokia making bank on their overly finicky 200Gpbs SFP and these guys "wires? meh"
That's a bold claim and pretty wrong.
One fun hack I wanted to make was a database function that would let me update database record on the GPU. One could implement the function as a kernel and have it update tens of thousands of records in one go.
The bottleneck is reading data from disk, to cpu / ram, to gpu.
Instead of there was a way to read data directly from disk, all records could be loaded in VRAM at once and data updated.
I think for a good number of use cases databases are prime candidates for parallel processing. Each record being processed by a kernel.
For gaming a significant bottleneck is moving texture and mesh data between disks, cpus, and the gpu.
Essentially, I think the main component of a PC could be the GPU while the CPU is there to execute block operations - if there is no way to “fuse” gpu cores as per my first paragraph.
I just think there’s a lot of performance left on the table otherwise.
For ml inference at edge for instance all you need is a GPU with a network card and extremely fast storage. All of which are available but bottlenecked by the cpu.
Technologies like GPUDirect are interesting in that vein. Theoretically you can DMA data directly from an NVMe drive on another machine, across an Infiniband link, and into VRAM without the data ever touching the CPU.
CPU independent DMA engines have existed since 1995 or earlier... DMA engines can typically be controlled by a queue of work, and can read/write to anything memory mapped. The work units themselves can be in memory mapped space of the GPU or network adapter, allowing those devices to queue up work too. And the DMA transfer can be triggered by any interrupt source too (initially intended for things like sound cards to have a tiny buffer be frequently refilled by the DMA engine without CPU intervention).
Basically, hardware has supported this forever. It is software that is lagging.
Consumer class GPUs don't have a lot of flexibility regarding VRAM size or type. GPUs have highly optimized, integrated memory controllers that are designed to use one or only a small number of distinct VRAM capacities and types. Enterprise class GPUs do have more options wrt VRAM, but it's still not a free-for-all like typical desktop/server RAM configurations.
I suppose one can imagine this changing in the future, but the memory controller complexity and costs (die area, power consumption, etc.) would need to be greatly increased.