HNHacker News
TopNewBestAskShowJobs

gabri200

1 karma · joined September 4, 2026

submissionscomments
gabri200··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
To get a good coding agentic system you need to use big context (Specs and conversation context can't be condensed every minute), so you need to use prefix caching, and the price for the hit cache tokens can't be the same that miss cache or the final price could be insane.
gabri200··on Qwen 3.8 27B available on Cerebras at 1500 tokens/s
I tried to get Qwen3.8 27B working properly under high concurrency, and while the quality level is spectacular for the size, the performance wasn't the best, even with MTP. Unless you have a very big infrastructure, it's difficult to run a dense model concurrently with high throughput.I suppose that's why almost all large models are now MoE. On the other hand, 1500 tok/s is an impressive speed, and that speed is very important for agent tasks, so a service like this instead of local infrastructure might make sense, although it also depends on your busines constraints.