The question is if the producers of these models were less incentivized to make them agreeable simply because most people don't like being spoken to like an idiot (or having their asks vetoed), how would they actually react? In the same way they exhibit emergence regarding their capabilities, perhaps "uncensored" in such a way they would convey some emergent behavior in terms of (at minimum) their "tone". Perhaps it would be interesting to see for examples if smarter models just by default became ruder or less friendly or aligned. Perhaps more aligned to things we would all generally agree on, but less agreeable to an individual ask. Perhaps sub agents would be less valuable for a whole suite of use cases if the agent itself was allowed to be more critical at the root. Idk. But I do not believe it is simply a matter of prompting alone.
Recent example: I asked Opus for some kind of reference data for MTBF in software. It took a few minutes to do "research", ended up providing zero data, and gave me a long essay on how DORA is superior and I should be using that instead. When I clarified that we're thinking of adding MTBF next to our existing DORA metrics, it decided I need to be thought about SLOs... if it was the first time, I may have continued clarifying but since it wasn't, I just gave up on Claude for this topic.