VLIW vector units optimized for AI (-ML suffix) with runtime reconfigurable network on a chip letting you optimize "streaming" the data between individual cores.
The way "AMD IPU" device was implemented, which embeds the system in some new AMD CPUs, the previous drivers (which were Linux/RTOS only) didn't work.
I was actually in the middle of reverse engineering how it was exposed in Phoenix (7940hs) APUs to write custom driver for Linux based on the one shipped for windows.
Which makes sense, because under Xilinx brand they have been selling accelerators and SoCs using AIE and AIE-ML cores for various use cases in embedded world, and XRT lets you program those with pretty much plain C++.
For some AI tasks, they have provided high-end wrappers for models that use ONNX (iirc) so those can be used immediately. This is not ROCm, and has essentially none of the ROCm bullshit in terms of support. I seem to recall there's some work to integrate parts of ROCm (opencl, HIP) with XRT.
Definitely going to build some kernels with this and test drive it ASAP.
https://ryzenai.docs.amd.com/en/latest/
https://xilinx.github.io/mlir-aie/
Looks like it's another NPU/TPU. Apparently it has a similar architecture as Versal.