It also sounds like from the paper transformers are formally proven to capable of it, but the training can't get the weights to make it capable of it, even though there should exist such weights, if I am reading them right.
It also sounds like from the paper transformers are formally proven to capable of it, but the training can't get the weights to make it capable of it, even though there should exist such weights, if I am reading them right.
In the paper I don't think they are looking for perfect, but they show which models can't seem to learn significantly beyond the size of examples they explicitly saw.
I missed it because it was spaced over a few lines and I was tired.
Pattern recognition can fail. Mine does, GPT's does.
I count further than I subitize, but from experience my mind may wander and I may lose track after 700 of a thing even if there's literally nothing else to focus on.
But I still didn't notice an error at depth 3.
You get the general rule (in linguistics known as competency) but can't flawlessly do it (performance). Transformers can't seem to get competence here.
on the ohter hand, thinking with the parenthesis takes language up to the level of functional-thinking. computable stuff is all about substitution of all things inside a parenthesis for a single thing. and this is the essence of all classic computation
Formally transformers should be capable of modeling it with the right weights, but they show good evidence those weights can't be learned with gradient descent. Whereas a neural turing machine or stack machine can learn them with gradient descent.
In standard automata theory, all inputs to automatons are finite.