Qwen-3.7-Plus is quite OK, good for subagent use. Way better then Sonnet.
Qwen-3.8-Max-Preview seems working just fine for me at the moment - I am playing with is right now but too early to say anything. At 10% of regular price it is a steal so far.
Qwen-3.7-Plus is quite OK, good for subagent use. Way better then Sonnet.
Qwen-3.8-Max-Preview seems working just fine for me at the moment - I am playing with is right now but too early to say anything. At 10% of regular price it is a steal so far.
I've been using https://gitlab.com/gabriel.chamon/orisun which is my own simplified methodology, for coding web apps in python and elixir and have been very successful using qwen3.6 27b Q4 locally with help of larger models for architecture, so I get very suspicious when people talk how useless larger models are. They are either using it for a domain that models don't perform well or just not using it right.
But it is how I feel and it feels like the right word for the job. Because as you say, good code projects start out with good decisions.
It's like when you see a CAD design with a sequence of features that exist only to fix problems caused by starting from the wrong principles or the wrong baseline.
Sure the resulting part may end up identical as a solid for that specific need, but it could have been done in a way that was more robust, simple, easier to understand and modify, and where the design doesn't break in an unexpected way due to a small change of an early measurement.
(CAD has made my instincts much more visible to me)
LLM will take stuff at face value. If you tell it it's good, it'll gladly abide. Similarly if say it's crap it'll so the complete opposite.
My latest experience was towards a C++ to Rust migration and it failed miserably because the existing codebase enforced patterns I didn't want.
It's important to do a first pass cleaning up and preparing code before unleashing agents onto it, similar to how Working Effectively with Legacy Code advocates approaching consolidated codebases. After that the agent will actually pick up the new pattern and start propagating that to the rest of the code.
> It's important to do a first pass cleaning up and preparing code before unleashing agents onto it
That’s a good idea in general, but the whole point of the rewrite was that Fable wasn’t capable enough of fixing the old codebase which we were using as a reference.
Another problem I had in another project was with having “sample” code purely written for testing being used as the architectural guidance, despite comments and repeated memories asking not to.
In the end, by the way, the solution was to simply blocking access to specific files, in the Claude settings.json.
But still Claude tried to cheat by running ‘cat file.cpp` a few times, so…
> At 10% of regular price it is a steal so far.
What price do you see?Here standard plan has been discounted to $18.00, from $25.00/month.
Such diametrically different ones.
I meant Opus 4.8 which is rather dumb and ineffective in coding harness, especially with higher thinking levels.
Opus 4.8 training works well for agentic work. Not for code harness.
EDIT:
```
stronger on coding and raw capability but can be more argumentative, verbose, and costly.
Reliability and instruction-following Many users say 4.6 felt more reliable and followed instructions better. "With 4.6, when I tell it something, it actually remembers the spirit of what I asked for and keeps applying it."
Others report 4.8 drifts from preferences and can be frustrating to control. "I still find myself getting frustrated when it ignores preferences and drifts from instructions"
Some people find 4.7/4.8 push back more and act more adversarial than 4.6. "The biggest complaint against 4.8 is that it is argumentative and "pushes back" constantly"
Coding quality and capability Several users praise 4.8’s coding strength and thoroughness. "4.8 is technically impressive, especially for coding"
Other reports say 4.6 could be better for certain coding workflows and breaks less. "4.6 still >> 4.8 for anyone else as well? Maybe I'm in the minority, but for my use cases Opus 4.6 is still better than"
Some recommend mixing models: use 4.8 for key tasks and 4.6 for general work to save tokens. "What I do is... use 4.8 for key moments, and for everything else 4.6"
Cost, speed and token behavior Users note 4.8 often uses more tokens and can feel slower because it “thinks” more. "4.8 is much more cautious, and as a result - slower. It checks everything, thinks for a long time etc."
```
[https://www.reddit.com/answers/601770d4-4059-478d-aa52-b445c...]
I'm finding the same with ChatGPT recently since the 5.6 release. Not as bad though, but sluggishness at times, harness churn (creating bugs and crashed), and occasional availability issues that cause me to downgrade to 5.5.
It's gotten to the point where I dread a new model release from these companies because it's guaranteed to be disruptive! I assume the pay per use API is less impacted.
I think you might both just be reading way too much into one-off random experiences that you've decided are evidence of significant and stable capability.
The model which everyone else raves about and is wildly successful with legions of programmers virtually demanding access while abandoning ChatGPT and Copilot in droves, is rather dumb?
Have you considered that it's more likely that you're doing something wrong?
In reality, they just aren't used to it.
Point is, sometimes people just have different experiences from you.
Interestingly with Fable vs GPT-5.6 I think they've lost their lead a bit. I'm finding Fable can't do certain work that 5.6 Sol Ultra can - especially when it comes to webpage design.
Grok 4.5 was fast but made mistakes that GPT/Fable just don't.
I'm curious to try Kimi.
Do you have a source for that? Codex went from 5 million users to 9 million users in the past few weeks since GPT 5.6 released. It was so popular that Claude was forced to extend Fable access by a week and then permanently for some plans.
It makes a huge difference if you're writing Javascript/HTML/CSS, Python, or C++/Rust.
Also the application type matters, e.g. user interfaces or scientific computing.
domain: typical web backend tier, mobile apps. not particularly complex, but requires OOP/architecture/system design.