More RAM, so less movement of the weights around to generate a token. Most of the speed limit on a LLM is bandwidth of getting the weights around. To a great extent, your token speed is approximately your (model size)/(effective bandwidth). If you need to shuffle the weights into VRAM from main RAM, you halve your speed (bandwidth used both to move into VRAM and out). If you need to pull the weights from disk, even worse.
While true, the benchmarks are not run on the Ryzen's NPU but the much stronger GPU.