They could take any conversation without suggested answers, truncate it to just before a user message, have the model predict suggested answers and then train it on the difference between predicted and actual answers, right?
They could take any conversation without suggested answers, truncate it to just before a user message, have the model predict suggested answers and then train it on the difference between predicted and actual answers, right?
If anything, showing the suggestion introduces unwanted bias.
This is similar to blog articles since 2023 or so onwards being less useful for model training since more and more of them are based on LLM output to varying degrees.
So you give the user a suggestion, and the user accepts -> good
You give the user a suggestion, and the user refuses and types something else -> bad (plus some supervisory training data)
The main performance enhancer in LLMs is getting high quality training data. So, first, any extra training data will help. Second this is training data that's directly relevant to their product, and thus higher quality than many other sources.
I'd believe any model provider is mining the shit out of every last customer interaction they can get, not just this.