So again, what is the actual difference you are imagining?
Or is it just that distributed X is fashionable?
Just because the swarm infrastructure hosting an LLM has higher latency across certain paths does not make it a swarm of LLMs.
Interesting, I haven't heard of that. Can you name examples?
If 80% of the processors in a cluster are running 'general LLM' and 20% are running 'math LLM' are they the same cluster? Could you host the cluster in a different data center? What if you want to test different math LLM modules out with the general intelligence?
In the case of the brain, while certain functional regions are highly specialized I would not consider them "a small separate brain". Functional regions are not sub-organs.
One LLM is limited, one obvious limitation is its context window. Using a swarm of LLMs that each do a little task can alleviate that.
We do it too and it's called delegation.
Edit: BTW, "swarm" is meaningless with LLMs. It can be the same instance, but prompted differently each time.
Better to limit his incompetence to one position.
So, perhaps, there aren't swarms yet just because there are easier ways to scale for now?
For the same reason one genius human does not suddenly need less support staff, they actually need more.
Edit: and why it isn’t here yet is because it’s new and hard.