> so their tok/s is a ceiling, not a true decode rate. The clear read: the GPT-5.6 tiers are the snappiest models here on short prompts (Luna answers in about a second), Qwen is absurdly cheap and fast, and DeepSeek and GLM are the slowpokes
You put in a lot of good work, and kudos for that, but man, reading paragraphs like these just puts me off of the entire piece.
Like…how hard would it have been really to type these two sentences by hand, in your own natural voice?