Tone: Cross platform audio tagger and metadata editor
github.com
github.com
I have a series of "mixtape" mp3 files and I'd like to generate a list of the songs within the one long file. I'm aware of musicbrainz, but is there anything that works on a single file with multiple songs? Perhaps my best bet is to generate 10 second long files sampled every two minutes, then run the fingerprinting on those individual files?
https://stackoverflow.com/questions/42507879/how-to-detect-t...
ffmpeg -i audio.mp3 -af silencedetect=n=-50dB:d=0.5 -f null - 2>&1
# -50dB is the threshold for detecting audio as a silence
# 0.5 means the silences must at least have a duration of 0.5 seconds
In your case I would try: ffmpeg -i audio.mp3 -af silencedetect=n=-30dB:d=2.5 -f null - 2>&1
That should give you a pretty accurate list of where the songs begin and end... that way, you can extract the songs and run fingerprinting.After this, you could use the ability of `tone` to add `chapters` to audio files (for mp3 it is id3v2 chapter addendum, see https://mutagen-specs.readthedocs.io/en/latest/id3/id3v2-cha...) to have at least marks for each track as title. MP3 chapters will not be recognized by most players, but at least it is within the specs.
If that is the case, track lists are using listed where the mix is posted, mixcloud, etc.
Otherwise, you could play a mixtape and put shazam in the background to see if it will find the songs (it can do that continiously). If it can, then you can easily automate file splitting using ffmpeg and powershell and grabbing a list using GUI/CLI automation via shazam desktop or somethig like https://audd.io.
Examples:
- Duration detection is inaccurate
- Tags `MovementName` and `MovementIndex` (for series) are not supported for some formats
- etc.
Another disadvantage is, that (afaik) `ffmpeg` needs to reprocess the file (audio wise) to write tags instead of just changing the metadata inplace.I ended up using mp3tag on windows and didnt see anything cross platform for free in a quick search last year
Real world data being: one on one interviews (no background noise), small groups of people chatting (lots of background noise), and specific audio recordings ( with varying British regional accents.
In all three instances whisper produced a more accurate transcription.
This is for personal use. The license of MMS is also restrictive so cannot he used for commercial uses while whisper can. Another key consideration when wondering what to choose. On the other hand, one can train MMS (so using own custom dataset) so for some projects it may be more suitable.
tone epub --format="markdown" --extract-sentences --one-file-per-chapter output-path/
As a result, you can use https://github.com/readbeyond/aeneas with the generated text / markdown files to create a json mapping file looking like this: {
"fragments": [
{
"begin": "0.000",
"children": [],
"end": "7.920",
"id": "f000001",
"language": "eng",
"lines": [
"This is the first sentence of the audio book."
]
}
}
Since aeneas is a bit inaccurate, I'm also working on an improvement with silence detection for these mapping files.If you are looking for something that is "ready to use", you could check out https://github.com/r4victor/syncabook or the according library https://github.com/r4victor/afaligner
If you have audio files, that are NOT audio books, the epub approach will not help you and the other comments are more helpful.