I remember seeing this popping up in discussions the first time, but never noticed any resolution (other than to train both sides). Has SOTA advanced?
It could be fun to try to make the model pre-learn a "reversal prior" that would cause a greater degree of generalization there, but I'm yet to see a published result like this. Let alone one that would demonstrate such a prior to be useful.