will catastrophic forgetting still occur if a fraction of the update sentences are the original training corpus?
is the real issue actually catastrophic forgetting or overfitting?
nothing prevents users from continuing the learning as they use a model