That 7 months of claude -> 16.5 months of claude.
Just do benchmarks yourself on the new model and decide if it is valuable for your usecase, even with the supposed nerfing.
Benchmarks are benchmarks. And you can ignore the data at your own risk.
If I’m using a calculator to verify my math, I don’t want to use a second calculator to verify the first one.
It was always random. This is no different than any other randomness that already exists in LLMS.
If you are concerned just do benchmarks and see if it is valuable for your usecase regardless.
Because of this there's a chain of trust between myself and the tools I rely on to do work. The people who create those tools see unpredictability as a problem, and that's the only reason I'm using them. I can't work on important systems with a vendor product like Claude Fable.
That being said there's plenty of work to do where it'd be amazing. This isn't an either/or situation.
I guess, given that, a pro tip would be to err toward sequential work rather than giving monster prompts. That constraint has got to degrade quality though.