ParentFull threadWanderPanda·Plus all the flexibility of CUDA is not needed for LLMs anywaysView on HN