curl `\
yt-dlp -j "https://www.youtube.com/watch?v=aeWyp2vXxqA" | \
jq -r '.automatic_captions.en[] | select(.ext=="json3") | .url'` curl `\
yt-dlp -j "https://www.youtube.com/watch?v=aeWyp2vXxqA" | \
jq -r '.automatic_captions.en[] | select(.ext=="json3") | .url'`To get youtube subs in .srt format this gave me some limited success:
yt-dlp --convert-subs=srt --write-auto-sub --write-sub --sub-lang "en,en-us,en-GB,automatic-caption-en" --skip-download "https://www.youtube.com/watch?v=1OfxlSG6q5Y"
Behind the scenes yt-dlp is downloading the subs in .vtt format than using ffmpeg to convert those to .srt. Depending on your situation the original .vtt format might be fine.Why not just use the whisper cli on yt-dlp CLI's output for videos with bad or no subtitles?
... | split_sentences | grep -viE '*vpn*'I'm sure there are many unix-y tools for this purpose, but I don't know of them. If you're looking for something that's installed everywhere, maybe a very big awk or sed regex with multiline wizardry could do the trick for most easy-to-parse latin languages and you'd just have to copypaste it around. It prolly becomes harder for regexes once you start working with right-to-left languages like Arabic, and languages with different ponctuation, so it might not be i18n-friendly.
Related Stackoverflow : https://stackoverflow.com/questions/33704443/python-regexp-s...