Full threadmwitiderrick·"What’s impressive is that the sparse fine-tuned LLM can achieve 7.7 tokens per second on a single core and 26.7 tokens per second on 4 cores of a cheap consumer AMD Ryzen CPU."View on HN