I also tried and failed here. We ran our own speech engine with a custom model- but it's extremely expensive as a cloud service, and incredibly tough to reach acceptably high accuracy in different environments. Adding NLP on top of error-prone transcripts will multiply the error rate and lead to all sorts of weird actions.
I really think on-device models like we see in Android's Live Caption tool are a major privacy boon, and they're starting to reach an acceptable level of performance in Google's case. The main pathway to better performance is loading ever-more-massive models into memory, which isn't feasible for mobile devices but could be done on people's laptops in a meeting.