> Is avoiding CF potentially just a matter of sheer scale ?
My intuition would be that you get more orthogonal directions to the gradient (of previous samples) if you have larger model.
My intuition would be that you get more orthogonal directions to the gradient (of previous samples) if you have larger model.
No comments yet.