For my style of coding (quick back-and-forths and corrections) it makes a big difference if a model comes back in 1-2 minutes compared to 5-10, and I am happy to pay a bit extra for that.
For my style of coding (quick back-and-forths and corrections) it makes a big difference if a model comes back in 1-2 minutes compared to 5-10, and I am happy to pay a bit extra for that.
As someone who only needs AI for a couple of tasks per day, I don't really care how much it costs, especially when subscriptions are subsidized. I want to filter by speed (eg, max task time < 1 min) and then choose the intelligence I need for the task. This will surface models like Gemma4 31B (xhigh) running on Cerebras and GPT Sol (med) fast mode. Using these models feels great and are affordable for infrequent tasks.
Go to https://artificialanalysis.ai/ and scroll down to the second graph under “Speed & Latency”.
I think this is the most import graph on their page. I wish they would let us filter by intelligence, or pass rate, and then see this graph. This is the tradeoff that actually matters, cost/token or tok/s can be very misleading (take glm-5.3-flash as an example).