A.k.a.: Too expensive for inference at fp16 and too dumb after quantisation, so it wouldn’t look good to release a step backwards.
GPT 4.5 was scrapped for similar reasons.
GPT 4.5 was scrapped for similar reasons.
If you signed up for a 16 core AWS server and they randomly kept changing it down to 8 cores, you'd sue them. How long until the same applies to these AI companies.
The amount of meddling they do to the harness, system prompt, model, quantization etc makes these products sometimes unbearable; you never know what you're going to get. A few more iterations of Qwen 27B and hopefully we won't have to deal with any of this malarky any more.