revised version by an another user
8 karma · joined January 1, 2024
I will imitate the good points of other players with ASR implementation.
I would like to improve accuracy by preserving context, but I haven't found a good way to do this at the moment.
If we are talking about the accuracy of the transcription, it is very good if you use a large model. At least the accuracy of whisper is far superior to Youtube's subtitle generation!
I will research if there is a good quality playback method locally, Thanks.