I ran into this recently by accident when writing a simple RL example. With two weight matrices to learn, the first weight matrix was given correct gradients, the last weight matrix was only supplied with partial information. Surprise, still works, and I only discovered the bug _after_ submitting to OpenAI's Gym with quite reasonable results. I've seen similar issues in the past such as accidentally leaving a part of the network frozen (i.e. it was randomly initialized and never changed) yet the model still happily went along with it.
This is good and bad. Bad in that it makes errors difficult to catch. Good in that, if you had a reason for freezing part of the network (maybe transfer learning etc) your model will learn to happily use it, even if that "information" is more akin to noise.
Regarding reproducibility, most papers I've gone to reproduce take far longer than expected and usually involve deducing / requesting additional information from the lead authors. Even minor issues, such as how the loss is calculated (loss / (batch * timestep) vs loss / batch) can confuse substantially and given they seem "insignificant" and that there are space constraints in papers, they are rarely written down.
Worst I have seen recently was a state of the art published result where the paper was accepted to a conference yet they didn't include a single hyperparameter for their best performing model - and no code. There is near zero ability to reproduce that given the authors spent a small nuclear reactor worth of compute performing grid search to get the optimal hyperparameters.
tldr There are reproducibility issues all the way up the stack, from gradient descent working against you to minor omissions in the papers to full fledged omissions that are still accepted by the community.