ParentFull threadMrScruff·LLM prompt processing and diffusion models are compute bound, while LLM token generation is memory bandwidth bound.View on HN