tone epub --format="markdown" --extract-sentences --one-file-per-chapter output-path/
As a result, you can use https://github.com/readbeyond/aeneas with the generated text / markdown files to create a json mapping file looking like this: {
"fragments": [
{
"begin": "0.000",
"children": [],
"end": "7.920",
"id": "f000001",
"language": "eng",
"lines": [
"This is the first sentence of the audio book."
]
}
}
Since aeneas is a bit inaccurate, I'm also working on an improvement with silence detection for these mapping files.If you are looking for something that is "ready to use", you could check out https://github.com/r4victor/syncabook or the according library https://github.com/r4victor/afaligner
If you have audio files, that are NOT audio books, the epub approach will not help you and the other comments are more helpful.
Real world data being: one on one interviews (no background noise), small groups of people chatting (lots of background noise), and specific audio recordings ( with varying British regional accents.
In all three instances whisper produced a more accurate transcription.
This is for personal use. The license of MMS is also restrictive so cannot he used for commercial uses while whisper can. Another key consideration when wondering what to choose. On the other hand, one can train MMS (so using own custom dataset) so for some projects it may be more suitable.