Tuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness.
I swear I spend more time telling Claude not to do things than telling it what to do.
I swear I spend more time telling Claude not to do things than telling it what to do.
Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.
But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?