Ah, the old "good, fast, or cheap; pick two" proves true once again.
Give it a few months.
It means something, because it an iterative workflow. If you're willing to burn tokens, it's possible for weaker models to implement tasks by incrementally improving drafts.
Some of us want fast food
Not if your use case needs speed. For one of my products I can't use an LLM that has a p99 of >700ms for TTFT.
If it could output 1k tokens per second but needed 4 seconds to produce the first batch of 4k, would that not be viable?