It's not just text, it's structured text. It often contains tables, sections, lots of things that can be drawn much more nicely.
And in juggler (and probably other harnesses too) the LLM can reply in HTML. So if I'm planning e.g. some UI changes, I might ask it "show me what this button will look like" and it'll reply with a picture of that thing in its response. No temp files to open or clean up, and fewer tokens burned