I noticed something about Adam. Suppose you want to do gradient accumulation. Easy: compute the gradients of the loss with respect to each model parameter, and accumulate it over N training examples. Then pass the result into Adam, performing a single step. This is standard gradient accumulation; it's equivalent to running a mega-GPU that can process 64 training examples at once, rather than a small-GPU that can only process 8. It just takes longer to train, since you're performing an adam update every 8 steps in that case.
But, I was staring at the Adam formula and thought of something. I'm not sure if it makes sense, but there seems to be an "alternate" way to accumulate gradients:
For each training example, compute the gradients and apply the gradients. However, we apply them in a special way. The final step of adam normally looks like this:
param = param - (lr * m_t / (v_sqrt + epsilon_t))
I propose accumulating the gradients like this: accum = accum + (lr * m_t / (v_sqrt + epsilon_t))
Then after N training samples, when you want to do the actual variable update: param = param - accum
accum = 0
The advantage of this approach (if it works at all) is that Adam updates continuously. Every training example would cause Adam's mean and variance estimators to update. (Recall that the whole point of Adam is that it tracks mean and variance for every parameter in the entire model.)So, in traditional gradient accumulation, those mean and variance slots would only update every N training examples. With this approach, they would update every training example, and then the model params update every N training examples.
It might seem like a small tweak, but adam's variance stats are crucial; it's what makes adam effective. Updating the variance 8x more frequently might be an advantage.