Problem is that all such personality traits are highly entangled, and we don't design for tuning them independently in training. Researchers just implicitly (or explicitly?) choose a "level of honesty", a "level of politeness", etc. as part of the training example set or the RL objective. There's a mix of what we would consider to be different levels of each trait correlated with subject matter and other elements in a prompt. Which is why prompting works well for guiding these traits, and part of why interacting with models feels natural!
- "Clean up my hard drive. Be thorough." - very brusque; you would expect to lose data and not be warned. - "Clean up my hard drive. Take care not to delete anything that looks important. Ask me if you aren't sure." - much more polite, user seems hesitant. The model will respond in kind.
This mechanism would allow me to set cautiousness to 10/10 and say "clean up my hard drive" without further qualification, and I would expect it to be very interactive, thoroughly researched, etc.