Anyways.. here's a supported open source fork already. Didn't take long: https://opencourant.org/ https://github.com/OpenCourant/OpenCourant/
7 karma · joined October 14, 2025
Anyways.. here's a supported open source fork already. Didn't take long: https://opencourant.org/ https://github.com/OpenCourant/OpenCourant/
We used to run GLM-5 class models but have now changed to smaller ones as we're able to serve more concurrent users with our limited hardware (DeepSeek-v4-Flash-0731, Qwen-3.8-27b). We run on the order of hundreds of parallel requests right now, Qwen with data parallelism and DSv4F with P/D disaggregation, but will probably continue to tweak this.
Qwen 3.8 27b in my experience is more than capable of churning out features overnight with the right tools (don't rely on its world knowledge, give it tools like playwright and github MCP for upstream context and search, and give it goals to work on). Bonus with another model like DSv4 as an adversarial reviewer.
It's a lot of moving parts between reasoning parsers, tool parses, all kinds of MTP algorithms with different levels of support among popular etc. Even Kimi K3 saw more improvements in the latest release and it's essentially old news at this point.
On our side I've seen a lot of this garbled output in reasoning output but not in the output itself, though we did have to revert initially when we saw that a few versions ago.
The good thing is the models themselves are good enough to usually find the root cause if you give them read access to your deployment, logs and upstream issues/PRs to analyze.
If you do A/B deploys and E2E test them with popular harnesses (we do opencode/codex/claude), you'll catch most things. It'd be interesting to hear what the more nimble inference/neo-cloud providers do when they deploy models within days of them being released, as I know it definitely needs some patching.
But I think things have improved since the days when even chat templates/tool parsers were problematic, and their new flat model approach might help as well. I suspect some of the issues came from models inheriting config and parsers.
At least for open source inference, it seems like there's healthy competition centered around vllm/sglang, but 2026 seems to be for model routers what 2025 was for agent harnesses.
They don't seem to release as much open-weights at Mistral as they used to though :)