> I'm just wondering why no one tried to RL a model on stuff like "less LOC" and "less overengineering"
Anthropic looked into this and the answer is because it makes the model more stupid
https://www.anthropic.com/research/evaluating-feature-steeri...