Aren't these still essentially completion models under the hood?
If so, my understanding for these preambles is that they need a seed to complete their answer.
If so, my understanding for these preambles is that they need a seed to complete their answer.
Also I wonder if it could be a side effect of all the supposed alignment efforts that go into training. If you train in a bunch of negative reinforcement samples where the model says something like “sorry I can’t do that” maybe it pushes the model to say things like “sure I’ll do that” in positive cases too?
Disclaimer that I am just yapping