Open Sourcing Active Question Reformulation with Reinforcement Learning
ai.googleblog.com
ai.googleblog.com
Nothing directly reward reformulations. But the global answer can be rewarded by user feedback. Yes this indirection still seems like an issue.
What I would do instead of this strategy would be to cluster extremely similar/other formulations of the same question by different users and then store for each frequent common question a list of reformulations (user generated). Of course the list would be based on the profile of the user.
I do not answer How similarity/identicality of user formulations would be determined by I have a couple of heuristics in mind.