"Notably, all DNNs face the issue of overfitting as they learn, which is when performance on one data set increases but the network's performance fails to generalize (often measured by the divergence of performance on training vs testing data sets)."
Not really. For example, "Gradient Methods Never Overfit On Separable Data" https://arxiv.org/abs/2007.00028