>AMD also claims its Strix Halo APUs can deliver 2.2x more tokens per second than the RTX 4090 when running the Llama 70B LLM (Large Language Model) at 1/6th the TDP (75W).
It's because of the bigger VRAM - 70B parameters don't fit into the 4090's 24GB.
Wait, this claims 50 TOPS. How is it faster than 4090 that does >300 TOPS?
More RAM, so less movement of the weights around to generate a token. Most of the speed limit on a LLM is bandwidth of getting the weights around. To a great extent, your token speed is approximately your (model size)/(effective bandwidth). If you need to shuffle the weights into VRAM from main RAM, you halve your speed (bandwidth used both to move into VRAM and out). If you need to pull the weights from disk, even worse.
While true, the benchmarks are not run on the Ryzen's NPU but the much stronger GPU.