AI is way beyond conventional LLM architecture now. It combines LLMs with search + RL. The traditional LLM architecture hit a wall around GPT-4o. Arc AGI evals show this.
We can, at best, approach a good set of weights, even in tiny neural networks.
Imagine if we found a way to calculate the exact optimal weights for a given loss function. I mean, there is an exact optimal solution, it exists, but we can't find it exactly, even for a neural network with just 50 parameters.
There is an optimal set of weights that minimizes the loss function for a given set of training data, but we cannot find it.
Granted, even if we could, it might just be overfitting.