I agree, the article doesn't argue the case very well. It's particularly tricky with the claim about babies, since most of the research on this (in general) deals with older children. (It's hard to study language acquisition in those who can't speak yet!)
I think you may be confusing emotional/gesture communication with visual speech cues. Because there really aren't that many visual cues for speech apparent with a mask on. Lipreading focuses on the lips, jawline, tongue, teeth, and nostrils (all covered up) along with the eyes and eyebrows and throat.
Anecdotally, I'm pretty much deaf, and I can lipread pretty well. I sometimes cannot even tell if someone is speaking or not behind a mask, let alone figure out what they're saying. I am poor at localizing sound. I'll hear a voice, or think I hear a voice, and I'm scanning all the masks trying to figure out which one is wiggling. With my residual hearing and lipreading I used to be able to pass as hearing in most contexts. These days I just do the "No, I'm deaf" up front because I just can't understand most people who are masked. The severity of it surprised me. I knew I relied on lipreading, I just didn't realize it was that much.
From that, it seems almost obvious to me that I would have had a much more language-poor environment in school under such conditions. Unfortunately I don't have much more than anecdote to rely on here; most deaf and HoH people I know have expressed similar sentiments.