AMD Upstreams AVX10_V2_AUX Support into LLVM Clang 24
phoronix.com
phoronix.com
It has a smells of 'we need to push back our IP deadlines' thingy.
Because, if I am not too mistaken, inference hardware would do all that and directly on the silicon, namely much faster and using much less energy.
Or is this to run quantized models on consumer CPUs??? Weird, because, if I do run models on my consumer system, I will want the full weights of the frontier open weight models.
This is one of the reasons IBM introduced an AI accelerator and added specific instructions to use it on their Telum II processor - so that limited inference can run in-process in transaction processing with latencies measured in clock cycles.
I think these instructions are for the case where inference is a minor part of the time it takes to process some data on that server - in that case the server will be more effective by having more, faster memory than having to go across a PCIe but to a GPU that will be idling most of the time, or going over the network to a dedicated system.