The important changes will be architectural, yes.
Three things I'm noticing:
1. We're reinventing structured programming in AI in fast forward. First it's a plain Markov chain. Then it's an "attention" directed acyclic graph. Then we realize we need loops. Then we realize we need to jump to different points in the loop. Then we realize it's useful to recursively call yourself or parts of yourself as a subroutine, parametrized with specific input. Etc.
2. Even before we fully realize this framework of thought into a model, I'm almost sure the model EVOLVES some of these structures during training. In the form of crude unrolled loops etc. Simply because it's inevitable for processing certain types of input data.
3. In order to preserve pragmatic outcomes, I'd bet the future is not one giant monolithic model for AI, but many medium-sized models, communicating in a meta network, like meta neurons, sending meta (high-level) messages to each other.
Essentially, we need to make neural networks more like a fractal. I have this rule of thumb that always works somehow: "no concept definition is complete, until it's made recursive". Neural networks will get there.