Fusing a 27B ternary LLM's whole decode step into one CUDA kerneltwitter.com3 points·Jr23_xd··0 commentsOpen articleSaveView on HN