Sounds like something a better model armed with ffmpeg would already be able to do. Run an analysis on loudness along the track, detect when speech begins/ends (with whisper) around loud segments, and compress or cut those parts out
250 karma · joined May 14, 2024
Also, I noticed that the gauge looks off (both colors and alignment) in e.g. iTerm, but renders much prettier in Zellij/Kitty. Any tips on that are appreciated!