1,378 karma · joined September 24, 2015
An LLM mostly deterministically (except parallel processing nondeterminism that can be mitigated) produces a probability distribution that can be sampled deterministically: just take the highest probability token or use beam search.
We'll have plenty of time for this, while living off UBI.
> but training does not produce such LLMs
If we are talking about fundamental limitations it should not be an empirical observation: "does not produce" (which is factually wrong, BTW). It should be a fundamental limitation: "can not produce in principle."
I still don't understand what you are talking about when you say "category mistake." I was talking about computational capabilities of LLMs with CoT that their training can exploit, not about Befunge-98.
What do you mean exactly? "Thinking is not an algorithm."?
The existing LLM training methods on the other hand give the results that are hard to distinguish from "thinking like people," judging by the end results.
LLMs with CoT are Turing-complete. So, theoretically, they can implement any kind of finitely describable algorithm (barring super-Turing computations).
Where are intermittent energy sources illegal?
I forget to add the obvious: some garbage gets through despite the mitigations. In the limit of pure garbage input, you'll get a model that internalized garbage generation. And you'd be better off throwing it away and starting over.
I guess my box labelled 'subconsciousness' is trying to say that low-level mechanisms that give rise to the observed cognitive phenomena might have nothing to do with neat boxes.
BTW, we seem to have Noisesphere instead of Noosphere.