ParentFull threadContinuityLab·Achieving 100 tokens/sec on consumer hardware for a massive model like this is an incredible engineering feat. Pushing high-throughput local inference forward is crucial for decentralized edge stacksView on HN