YouTube Transcript – read YouTube videos
youtubetranscript.com
youtubetranscript.com
curl `\
yt-dlp -j "https://www.youtube.com/watch?v=aeWyp2vXxqA" | \
jq -r '.automatic_captions.en[] | select(.ext=="json3") | .url'`Why not just use the whisper cli on yt-dlp CLI's output for videos with bad or no subtitles?
To get youtube subs in .srt format this gave me some limited success:
yt-dlp --convert-subs=srt --write-auto-sub --write-sub --sub-lang "en,en-us,en-GB,automatic-caption-en" --skip-download "https://www.youtube.com/watch?v=1OfxlSG6q5Y"
Behind the scenes yt-dlp is downloading the subs in .vtt format than using ffmpeg to convert those to .srt. Depending on your situation the original .vtt format might be fine. ... | split_sentences | grep -viE '*vpn*'I'm sure there are many unix-y tools for this purpose, but I don't know of them. If you're looking for something that's installed everywhere, maybe a very big awk or sed regex with multiline wizardry could do the trick for most easy-to-parse latin languages and you'd just have to copypaste it around. It prolly becomes harder for regexes once you start working with right-to-left languages like Arabic, and languages with different ponctuation, so it might not be i18n-friendly.
Related Stackoverflow : https://stackoverflow.com/questions/33704443/python-regexp-s...
The extension I made to export the transcript was based on this YouTube functionality, I should update the instructions now.
https://chrome.google.com/webstore/detail/youtube2anki/boebb...
[1]: Searching YouTube videos with transcript - https://needgap.com/problems/88-searching-youtube-videos-wit....
I had done something like this a couple of years ago for some specific set of videos (e.g. https://shreevatsa.net/tex/program/videos/s10/ — compare with https://youtubetranscript.com/?v=_0Cv1G_s4gQ for the same video), but never got around to making it general; glad someone has done it. It takes just a few lines of Javascript, using the Youtube API, to do this i.e. keeping the video and text in sync (just view source on either page to see the JS at the bottom).
Something like this can also help with audio recordings (generating the alignment automatically is called "forced alignment" and there are tools like "aeneas" for this). In case anyone's interested or wants to help (for Sanskrit texts): see https://github.com/shreevatsa/web-align-audio-text deployed at https://shreevatsa.net/ramayana/sarga/ and better version at https://github.com/avinashvarna/audio_alignment deployed at https://avinashvarna.github.io/audio_alignment/
We're doing forced alignment with audio recordings i.e. podcasts. Here's an example from a test client: https://www.withfanfare.com/p/seldon-crisis/search-by-the-fo...
Grateful for any feedback you might have.
Also, if you run a knowledge-dense podcast, or know somebody that does, I would love to talk to you/them. I'm for example considering to link places, people and things in a transcript.
You're doing cool work, good luck with your service! It looks very useful and if someone I know is running a podcast I'll recommend it for sure; it looks really well-done and polished (at least the example you shared).
We know the start and end timestamps on a word level and we know the current player time. So all we do is set a CSS class on the words that are currently playing. We not only highlight the current word, but also the words that are close in time. This is what generates the effect.
We just released autoscroll, which is something that was requested often.
Excuse the question, but "forced alignment" is when you don't have timestamps, like in webvtt?
Yes, exactly. We do forced alignment when you edit your transcript. The new words don't have any timestamps, so we need to align them. For short sections we use interpolation. If we need align whole sections we use Gentle[^1].
https://github.com/ggerganov/whisper.cpp/blob/master/example...
for my purposes I changed the output from subtitles to txt (so I could pipe the result into chatgpt)
Tell us more :)
See my comment on how to just download the subtitles YouTube provides with `yt-dlp` here: https://news.ycombinator.com/item?id=34040342
I'd imagine it's very similar for others. Often a company will pursue a violation if only to be consistent in showing the courts they actively defend their copy right.
They may not like it, but they don't have the power to disallow you from using their name to refer to them. That's allowed.
The law is on their side.
No, they don't.
> It's copyright.
No, it isn't. It's trademark. There is no such concept as copyright in a single word.
> The law is on their side.
Again, no, it isn't.
What was the point of your comment? Why talk if you're not worried about whether what you're saying is true or false?
Otherwise: https://wordpressfoundation.org/trademark-policy/
FFS Relax. It was an oversight on my part. It's minor at best in larger scheme of things. Maybe you're having a tough day? God bless you. But this is nothing to go to the mat over.
You could have just said, "Oh. Maybe you're confusing trademark and copyright?". The last thing the world and HN needs is another high-strung belligerent asshole. That doesn't help anyone, or anything, sans your ego. It's not a good look.
That is your most minor mistake. The much larger one is that you are claiming that the law supports the trademark policy you link to. It doesn't. Everyone remains free to use a trademark belonging to wordpress in order to refer to wordpress.
You are not free to use any such mark - the WordPress Foundation would be on very solid ground in telling you to stop using their graphical logo. But they can't stop you from using "WordPress" in the name of your for-profit company, which you'll note is quite explicitly contrary to their stated policy.
> is a legal doctrine that provides an affirmative defense to trademark infringement as enunciated by the United States Ninth Circuit, by which a person may use the trademark of another as a reference to describe the other product, or to compare it to their own.
And also save yourself time when the creator teases that they provide the info, but it turns out they don't, they're just trying to get views.
No snark intended, but i just gave up with the dross. And even some of them, of late, are getting a bit crafty. But, creators get one chance from me now - give me decent content, or even with the fancy chapters, you're not getting my eyeballs past two minutes. What I have found is that leaving the decent stuff on, what auto-plays after is 'generally' of similar quality. A quick set of back-buttoning and bookmarking has fairly often got me some interesting results.
A search engine could, narrow in on the few sentences AV in the video that it thinks correspond to what I was searching for, and summarize that, and also link me to the AV start timepoint in case I also want to watch the video.
This might change the economics of some YouTube video content creation.
Is this what you mean? (Pardon if I'm not familiar with the latest Google Search features; I've mostly been using DDG lately, so don't have occasion to see all the features that exhibit only occasionally.)
Similarly Google incentivizes longer webpages, so now we have recipes that start with a novella about grandma’s cooking before showing the actual recipe.
It used to be nice to see a video’s thumbs up to thumbs down ratio to know if you’ve been click baited or not before watching the whole video. But that signal has been removed now too.
recipe reader
I want to
dismiss cookies, have a video ad follow me down the page, and read why this cake conjures up memories of the author's childhood, before reaching the actual recipe
so that
I feel connected to the author, before fully committing to mixing ingredients
Including the part about declining more cookies offered (to save room for grandma's lasagna).
Or maybe just repeat the same content a few times.
Helped a lot with take home exams.
"Streaming your phone screen to a TV is something many people at some point want to do. But seeing the picture of your smartphone on a television or other device can be a daunting task. Here we list multiple ways how you can achieve the goal of sharing a mobile screen of an Android or an iOS device on a different device. It works for any manufacturer, like Samsung Smart TVs or Apple TV."
(goes on for another 15 paragrpahs before presenting a non-solution)
It's hard to put in words the amount of hate I feel for the authors of such pages. I'm desperately looking forward to the day Google comes up with better models for detecting useful content and those trash piles can burn in hell.
I’m just not convinced the preambles are needed.
If you go to the homepage with clear cookies it's just endless amounts of utterly dogshit cookie cutter content. Same clickbait thumbnails with a person pulling an idiotic expression. Even the videos masquerading as educational are entertainment at best. If I had kids I'd do everything in my power to keep them away from YouTube.
Google Search actually indexes transcripts of a video and shows you some YouTube results based on that even though the title/description of the video doesn't match the search query.
I’m considering a master keyword list to index against any text that comes in.
Any addition that is on the "view" layer (the generated HTML) is very easy to add, just needs to go into the template file, at some point I might tweak that area but currently have no outstanding idea/requirement. The rest (i.e. the bulk of the code) is just a very bare-bones parser for captions that should be pretty stable and need no additions (crossing fingers here).
I’ve seen vosk used on device and it’s decently quick too on a recent Apple chip.
To make transcripts easier to access might create more problems than it solves.
Granted I can't make a bullet-proof argument; there's no clear way to quantify that ratio.
My one gripe with Youtube's own transcript box is that it is too narrow, so it is a shame that a website designed to specifically make the transcripts more readable also displays the transcripts in a narrow box.
This will work as long as YouTube doesn't change anything. And since when has YouTube changed anything?
https://github.com/ricklamers/ChatGPT-YouTube-summarizer
All jokes aside, I love the automated transcripts from YouTube. Videos are just so inefficient to consume as a format.
> No transcripts were found for any of the requested language codes: ('en',) For this video ([...]) transcripts are available in the following languages: [...]
It even knows what language is available, so why no dump that instead?
Edit: after reading other comments it seems this may be using an undocumented api to retrieve the data.
https://demo.robomotion.io/designer/shared/6j984jBCQqYVBCaQk...
#!/bin/sh
ttml2srt()
{
x=$(echo x|tr x '\34');
tr -d '\34' \
|sed -n "/<p begin/{
s/<p begin=\"//;
s/\" end=\"/ --> /;
s/\" style=\"s2\">/$x/;
s#</p>##p;}" \
|sed = \
|sed "/$x/!s/^/$x/" \
|tr '\34' '\12' \
|sed '/[ ]-->[ ]/s/\./,/g'
}
read x;
case $x in
https://www.youtube.com/watch?v=??????????\
|https://www.youtube.com/watch?v=???????????\
|https://youtu.be/???????????\
|https://youtu.be/??????????)
f=${x##*=};f=${f##*/};case ${#f} in 10|11)
curl -4o $f.mp4 $x
video=$(tr \{ '\12' < $f.mp4|sed -n "/itag=22/{s/u0026/\&/g;s/%3D/=/g;s/%2C/,/g;s/%26/\&/g;s/.*url\":\"//;s/\".*//p;}"|tr -d '\134')
test $video||exit
ttml=$(tr \{ '\12' < $f.mp4 |sed -n '/timedtext/{s/u0026/\&/g;s/.*:\"//;s/\".*//;s/$/\&fmt=ttml/p;q;}'|tr -d '\134')
test $ttml||exit
curl -s4 $ttml|ttml2srt > $f.srt
exec ffmpeg -v quiet -y -i $video -vf subtitles=$f.srt $f.mp4
esac
esac
exit
The script above, "1.sh", can be used as follows. echo https://www.youtube.com/watch?v=aeWyp2vXxqA | 1.sh
It will download the captions as .srt and then "hardsub" them into the video as it dowloads the .mp4. NB. This is slow YouTube downloading without using yt-dl/yt-dlp. Obviously, it will not work with commercial videos.The .srt file is saved as [YouTube ID].srt and the video as [YouTube ID].mp4, where [YouTube ID] is a 10 or 11 ASCII character string.
The video format is itag=22, i.e., mp4/720p. Not all videos will have 22 of course. I usually try itag=18, mp4/360p, if 22 is not available. Change the format to whatever is preferred.
Looking around for a .ttml/.vtt/.srv[1-3] to srt converter I found solutions that required installing Python or some other large scripting language. On GitHub I found a project called "astisub" that will convert from ttml or vtt to srt. It is a 3.8M Go binary. I wrote a shell function instead.
We want to partner with you on a topl that autogenerates clips of any video based on the topic start and end
Any chance timestamps could be added?
Often enough I'd rather read than watch. Reading in faster. Having corresponding visuals would be a big plus.
Why would you reach for a cluster of machines working in parallel, when you could retrieve the already auto-created transcript from YouTube servers?
Also, other comments have pointed out that the transcripts are identical with the ones created by YouTube, which would be unlikely to happen if this service was creating transcripts of their own.
Summary remains the elusive hard part...both to do and to serve at scale. But we get closer by the day.