You can also run massive amount of LLMs in parallel.
There might be a limit to a normal LLM but not to theo everall system.
You can also run massive amount of LLMs in parallel.
There might be a limit to a normal LLM but not to theo everall system.
One instance of an LLM is the same as another instance, so while you may get more out of it by stacking more of them, I strongly suspect it falls victim to diminishing returns. 100 instances of the same LLM may converge on the same result as 10.
It's worth a shot at least, as a microservices architect I have a bias that we aren't networking these enough, a single main agent session orchestrating multiple subagents is different from multiple main agent sessions with their own subagents coordinating with each other.
Although I'm also not sure about just how much better models can really get with this technique. Ultimately you're still getting the same model with the same training data, which are the important parts. Asking it to pretend to be something feels like it would just put a color filter in front of the conclusion the model has already predicted, or maybe alter the path to the conclusion slightly or pick a less likely answer that it still could've provided normally.
FWIW I mean if I have an AGENTS.md that encodes my software heuristics (use an interface in situations like X, here's how we name variables, etc.) it generates far cleaner code than if I don't.
Edit- mostly pointing out that stacking 10 base models vs. 10 models with sufficiently different base context isn't necessarily the same attention routing. I suppose I was thinking about tasks that don't have a concrete single answer.
Not in the highly verifiable domains. There you can take it from say 80-90% maj@x to 99% pass@n. Math, some parts of programming and cybersec are examples of highly verifiable domains. (e.g. if you're searching for a linux LPE, that's expensive to search but easy/cheap to verify - just have a token in /root and have the model retrieve that token)
LLMs scale well in almost all dimensions. Context window (working memory) can be a bottleneck but for humans you can’t scale it at all.
Bigger limit and no limit are very different.