I looked into this problem a while back and haven’t looked at since.
The base ai model sounded like whisper ai from meta. Did you train the voice yourself or is it one of defaults?
I am always curious as to what copyright issues products like this run into. Also whats the stack like to build something like this?