That's not entirely true. As someone who has a pretty long background in this area, works with LLMs/Diffusion models everyday, and generally thinks there is a bit too much hype (but also a lot of potential): there is a lot that we don't really understand about how these models behave and why.
For starters, this article discusses how we need to change the architecture for solving XOR. That's something we do understand quite well. However what we really don't understand is why architectures like transformers work so well. From an engineering standpoint they make sense because the models look like they're doing something we want, and it makes intuitive sense that they work.
But from a theoretical standpoint it's not known why we really need all these fancy architectures (rather than just using a bunch of layers that should also be able to "figure out" what the network needs to do). All of our success has boiled down to "hey, let's try this and see if the model will learn better/faster"
Similarly, from a mathematical perspective, we do have some intuition around the reality that all NNs are basically doing some highly non-linear, complex transformation on to some latent surface where the problem is linearly separable. That gives us a sense of why these models probably can't learn "truth" (unless you do believe there exists a latent space of what is true and what is not, which is pretty radical). But if you start asking many more questions about how this works, or why the model would choose one representation of the other internally, we don't really know.
Nearly all of our progress in deep learning that last decade has been basically hacking around and applying larger and larger amounts of data and compute resources. But at the end of the day even the best in the field don't really understand exactly what's happening.
Following from this: If you really want to understand these tools better, start playing with them and trying to build cool things. A deep understanding of the fundamentals is not much more useful for success with LLMs and Diffusion models is that knowing how to efficiently implement b-trees helps build a cool product with a database back end.