on one hand , yeah : a model being more aggressive towards exploration of the decision space is usually a good thing.
on the other hand : I think that it's up to the operator to set rigid test criteria to make these things actually work well in a repeatable fashion.
so in other words, i'm glad claude is doing a better job for the way you're prompting the thing, but as a spectator from afar these kind of operator complaints usually spring up from the use of weak, ambiguous, under-considered prompts.
similar with the parents' complaint; there should have been a testing criteria for legible output that got immediately flagged or failed by the larger model.
it's a very hard sell for me to think that " a a a a a a a " is accepted as a valid language output by any near-SOTA-large model without some real coercion.