I wonder if it's less accurate for specific people, but more accurate generally. In other words speech recognition was more accurate in the beginning on english-speaking men, and maybe now it's better for women, kids, people with accents, other languages, etc...
They train this on examples of speaking and maybe it's more broad now.