Any architecture that could compete with LLMs is
going to reshape the ML landscape that is now
dominated by GPU-heavy algorithms like transformer attention.
Its likely there is something more fundamental that
could work without much training, like
finding exactly what training process changes and\
forming architecture around optimizing that.
I don't think people realize how much influence
'dumb stochastic parrots' have now, and the whole
progress in identifying what exactly happens inside
these networks is still a mystery: there hints
that by sheer size they forming structures capable
of cognition but only as side effect(finding out what
exactly is going there and replicating in simpler terms
would yield massive speed-ups).