These "pure speech" models could really benefit from being coupled to a large language model like ChatGPT.
YouTube live transcriptions are terrible, because they get confused by homonyms and can't follow the context in a sentence.
In the same manner that Dall-E joined an LLM to an image generator, they ought to train a combined speech model + LLM so that the uncertainties in the speech model output is disambiguated by the LLM.