Short answer is no, certainly to the "perfect" part.
The core problem in ML is generalization; simply put - how well does your approach work with new data it hasn't seen before. Think of it this way: there is a large set of all the potential inputs you could see, and you only get to see small subset when you are training; what do you do so your general performance is best? Which of course you can't actually know but you can try and estimate.
There are two issues that can give you a lot of trouble here. The first is overfitting (you'll do much better on the training set than "in real life"), the second is bias in in your training samples. Data augmentation (what you are talking about) is one approach to reduce parts of the former effect, and done correctly it can help.
Take a simple example, imagine we were trying to recognize simple geometric shapes on images of a page - you want to find triangles, rectangles, ellipses, etc. I only give you a small set images, say 10s of shapes total.
Now you suspect that "in the wild" you can have triangles at all sorts of rotations, and sizes, but I've only given you a few examples. So you want your algorithm to learn the shapes, but not the sizes or orientations. If you just train on these, it may not recognize a triangle that is just 2x as big as any it has seen, or rotated 20 degrees left from one it has seen, etc.
One approach would be to try and find a rotation and/or scale invariant representation for your inputs - if you "know" that shouldn't matter you've now removed it from the problem. This can be hard or even mathematically impossible to do, depending on the problem space (e.g. there is no rotation and scale invariant manifold for photographic images). So another way you can approach it is empirically; to take the examples I gave you, and generate new examples in different poses and scales. You feed this into your training and should get a much more robust result, one that doesn't hew too closely to the training set (i.e. less over training).
So this sounds great, right? What could go wrong? There are a few issues. One is you are now enforcing things outside what you learn from the data, so if you are wrong you will make things worse.
More subtly when you do this you can tend to amplify any of the sampling biases you had originally. Imagine that I never gave you an equilateral triangle in the training set. It's quite plausible that by generating millions of inputs from a few examples, this category gets pushed closer to something symmetric, like a sphere , say.
Another issue that can be subtle is that the manipulations you are doing for data augmentation can easily introduce new things to the data that you don't see, and your training can pick that up. Consider, for example, rotations of these shapes. I told you we were doing this from images, i.e. discretely sampled grids. This means that other than certain symmetric rotations and flips, you can't do this without resampling. And you can't resample without smoothing. So if you take a dozen or so "crisp" examples and turn them into 10s of thousands of "smoothed" examples, what exactly are you teaching your model. I'm also waving my hands hear about how you are extracting "shapes" from "background" and in a NN context, what your inputs actually look like... but you can introduce issues here also.
There are lots of trade offs here. It's a useful technique, but unsurprisingly isn't a silver bullet.