"Top down" never really works for complex systems, the economy being an obvious example. But we tend to ignore that when thinking about neural networks.
"Top down" never really works for complex systems, the economy being an obvious example. But we tend to ignore that when thinking about neural networks.
I took a deep learning class before things really took off [0], and what we were taught was that a lot of deep learning research was just trying different structures to see what worked. Even the people who were good at it didn't base their architectures on strong theories for how and why things would work, they just had developed an intuition through lots of trial and error.
All the explanations for what the various structures in a deep learning model are doing are post hoc—they're observations made of the behavior of the systems after we build them.
[0] Edit: Not just a random deep learning class either. The TA for that class created AI Dungeon that semester. The professor came in one day with a story of how they accidentally racked up tens of thousands of dollars in Google Cloud bills in 24 hours when that went viral.
That would be slightly annoying, but they did this to then justify points about how the resulting model is "emergent". The research process, and the resulting output, can and (my main point) almost always do, differ as to whether they are truly emergent.
I wish the industry would be more transparent about where all these brilliant ideas come from (spoiler alert: academics from the 1930s-70s, not "genius founders").
I wouldn't characterize it like that ... The brain has a specific learning architecture / dynamics that has been created via evolution under selection pressure to learn (i.e predict outcomes) better.
As a product of evolution I wouldn't want to call it "architected" or a top-down design, but for the time being it's the only example we have of such a successful learning system so it would make sense to copy what evolution has done, which means copying how it works on all levels.
A simpler example is convolutional neural nets for vision which were explicitly designed to simplistically mimic some of the behavior of our visual system with it's multi-layer (V1, V2, etc) architecture and local learning rules. Sure there's emergent behavior occurring when data is fed into such a system (both brain's visual system or CNN), such as the pattern detectors we see emerge at lowest level, but this emergent behavior is a result of the overall architecture.
[1]: https://en.wikipedia.org/wiki/Friedrich_Hayek#Economic_calcu...
[2]: https://en.wikipedia.org/wiki/Catallaxy
[3]: https://citeseerx.ist.psu.edu/doc_view/pid/2688969848b753368...
https://nonint.com/2023/06/10/the-it-in-ai-models-is-the-dat...
Even then, the architectures we use are essentially stumbled on. This is alchemy, not modern chemistry.
That said, I don't think that LLMs are yet close to fully utilizing the training data. e.g. they are prone to hallucinating due to not thinking ahead and backing themselves into a corner where they have to say something. One obvious improvement for that one is more forward looking prediction (i.e. "thinking ahead" - engage brain before opening mouth), which for LLMs can be addressed by tree of thought rollouts and RL learning.
So, while architecture is not important if all architectures are equally powerful (in which case you'll learn what's available to learn), it certainly does matter if not all are fully up to the job, as it would appear none currently are.
Yet all four of these examples use the same model architecture (transformers).
Without tweaking anything (so not RWKV), you could train a GPT level RNN...if you had the compute to burn.
We use Transformers today in large part because they got rid of recursion and in effect could massively parallelize compute.
The eye is part of the brain for instance. The details are emergent but there is most definitely strong top down architecture encoded in DNA