do these llms get trained with something like a credibility weight on the training data? that was it seems they did in this paper, just manually curated that
do these llms get trained with something like a credibility weight on the training data? that was it seems they did in this paper, just manually curated that
Think common use cases. A lot of users are students, do I want it to write an essay like a linguist? Or solve my homework using the better but more advanced techniques and style?
You do want the underlying model to be capable of the advanced answers, since if it is, it can be used to supply simple answers. You can't make that work the other way around in the same way.
If you have input data that includes a high rated reddit eli5 question, the content of that answer might be hard to verify, the style and way its delivered would be ideal to keep around in the training data. on the other side, technical in-depth answers have content that is worth keeping around, the style of its delivery would be very specific.
Keeping the entire internet around in your training data would still give you access to all these types of delivery still. hope that makes sense.