I hate this idea that our voice has to be shipped somewhere to be processed. I remember a lot of the speech-to-text tools in the early 2000s weren't all that great (they needed a lot of training), but why haven't we been able to advance on-device processing? Why is everything done in "the cloud."
So the only way to semi-accurately do voice recognition is to source algorithms that re-train off of millions of people? We have processors in our desktops and laptops that dwarf that compute power by leaps and bounds. We should be looking to Star Trek TNG level voice processing, on each individual device, without some central mainframe.
But marketing, advertising revenue, data mining, free (as in beer) software that pumps your data like an oil rig, efficiency in data centre (cloud) design .. all these factors have led to these powerful little Intel/ARM/Ryzen chips to be nothing more than thin clients when they're not playing games.
If Mozilla really wanted to make something amazing and in the spirit of Firefox, give us an experiment where voice processing is done on our devices. Even if it meant I needed to download a 230GB data set, I'd gladly do it, if it could remotely help in getting away from these data silos.