What brings to my attention in this article is the section named "Cold Start", where it generates questions based on a provided context.
I think it is a good way to cheaply generate an Q&A dataset that can later be used to finetune a model.
But the problem is that it generates some questions and answers of bad quality. All generated examples have issues:
- "What is the context discussing about?" - which context?
- "The context does not provide information on what Ray Tune is." - Not an answer
- "The context does not provide information on what external library integrations are." - same as before
I could only think of manual review to remove these noise questions. Any ideas on how to improve this QA generation? I've tried it before, but with paltry results.