The gradient we want is the gradient with respect to the process which generated the dataset. The gradient we get is an estimate based on only a handful of samples from that process at a time. The analogy holds up fine.
If we were trying to describe stochastic gradient descent, this would be relevant, but we're talking about backprop (which does often use just a batch, but that's not inherent). And there is nothing about backprop that makes the kangaroo more blind than in any other form of gradient-based optimisation.
I would say it's more a commentary on the fact the the gradient is effectively based on L2 distance in parameter space, and so can be a bad/inefficent direction to move in even if you have access to the full gradient. Hence the motivation for momentum and second-order optimization.