Continuous learningin current models will lead to catastrophic forgetting.
is the real issue actually catastrophic forgetting or overfitting?
nothing prevents users from continuing the learning as they use a model
In RL it can be that you are not getting meaningful data anymore because you are 'too good' and dont get anymore the "this is a bad answer" signal so you can't estimate the gradient.