There’s this sort of local optima all these models pick which is instantly recognizable and hardly digestible for human consumption. Perhaps too much internal RL against benchmarks during chain of thought? The older non reasoning models had they’re own problems but at this point I can’t get an LLM to summarize data for humans, which is troubling.
As much as I’d like to read this piece it follows suit and I don’t have the patience to try and extract anything valuable from it.