Very impressive speed. With a context window of 40K however, usability is limited.
> Cerebras Systemstoday [sic] announced the launch of Qwen3-235B with full 131K context support on its inference cloud platform
Then later:
> Cline users can now access Cerebras Qwen models directly within the editor—starting with Qwen3-32B at 64K contexton the free tier. This rollout will expand to include Qwen3-235B with 131K context
Not sure where you get the 40K number from.
Also this model https://huggingface.co/Qwen/Qwen3-235B-A22B
Is native 32k. So the 64k and 131k use ROPE that is not the best for effective context.
While https://qwenlm.github.io/blog/qwen3-coder/ it's 256k native https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct.