That substitution is the main difference between SGD and RWP.
It’s like describing bubble sort when you meant to describe quick sort. Would not fly on an ML 101 exam, or in an ML job interview.
It’s like describing bubble sort when you meant to describe quick sort. Would not fly on an ML 101 exam, or in an ML job interview.
The meaning of the gradient is perfectly adequately described by the author. They weren’t describing an algorithm for computing it.