In the paper I don't think they are looking for perfect, but they show which models can't seem to learn significantly beyond the size of examples they explicitly saw.
I missed it because it was spaced over a few lines and I was tired.
Pattern recognition can fail. Mine does, GPT's does.
I count further than I subitize, but from experience my mind may wander and I may lose track after 700 of a thing even if there's literally nothing else to focus on.
But I still didn't notice an error at depth 3.
You get the general rule (in linguistics known as competency) but can't flawlessly do it (performance). Transformers can't seem to get competence here.
on the ohter hand, thinking with the parenthesis takes language up to the level of functional-thinking. computable stuff is all about substitution of all things inside a parenthesis for a single thing. and this is the essence of all classic computation