Then we had RNN, which have the output fed back into the input. This gave the networks some form of memory.
Then we had Transformers, which are basically parallel processors. I.e generate 3 outputs in parallel, multiply them together. This was basically just a better form of compression applicable to everything.
The general trend here is that someone discovers some architecture that works out nicely and then everyone builds something around it. This is probably going to be the future. Google has some neat things with automated robotics, OpenAi has their A* stuff thats supposed to be "accurate" instead of probabilistic.
Then there is the hardware piece, which I know much less about, but hoping companies like Tinycorp or Tenstorrent give us a way to reliably run something like GPT3 full parameter model at home.