I don't have time to look into methodology in great detail, but their model seems to be flawed?
CDQ is designed to correlate with age[1] and it's expected that the most recently born children have lower scores. Just treating age as a model covariate seems wrong? Shouldn't they do some kind of propensity score matching (match equally old childs with comparable maternal education, weight, etc.)? instead?
I also really doubt that a verbal development assessment already makes sense for children born after March 2020 (or you would have to limit the sample to children born in March/April 2020 and assessed just now)?
Would be nice if someone who is more experienced with this data + methodology can comment on that...