ParentFull threadricardobeat·Most LLMs cannot run efficiently on current NPUs (except for prefill stage), the hardware was built for a different kind of ML workload.View on HN