It’s faster at AI than an Nvidia RTX4090, because 96GB of the 128GB can be allocated to the GPU memory space. This means it’s doesn’t have the same swapping/memory thrashing that a discrete GPU experiences when processing large models.
16 CPU cores and 40 GPU compute units sounds pretty parallel to me.
Doesn’t that fit the bill?