TexTube: Chat with any YouTube video transcript in ChatGPT fast
chatgpt.com
chatgpt.com
I'm still doing the last testing of the site, but might as well share it here since it's so relevant:
https://youtubetranscriptoptimizer.com/
There might still be a few rough edges, so keep that in mind!
In the case of a document only workflow, we generally want to stick to what's in the document very closely, and just extract the text accurately using OCR if needed (or extract it directly in case we don't need OCR) and then reformat it into nice looking markdown-- but without changing the actual content itself, just its appearance. When we've turned the original document into nice looking markdown, we can then use this to generate the quizzes and perhaps other related outputs (e.g, Anki cards, Powerpoint-type presentation slides, etc.).
Because of that fundamental difference in approach, I decided to separate it into two different apps. But I'm planning on using much of the same UI and other backend structure. The document centric app also seems like it has a broader base of potential users (like teachers-- there are a lot of teachers out there, way more than there are YouTube content creators). I started with the YouTube app because my wife makes YouTube videos about music theory and I wanted to make something that at least she would actually want to use!
The vodcasts that most need transcription are long form. After the "don't make me do math" pricing, you do have a table of minutes, up to 60, so for a typical, say, ContraPoints vodcast episode, you multiply by 3, and find out that could cost $30 to turn into the optimized transcript. (Which the creator might well pay for if they value their time, but viewers might not.)
Does your tool work on 3 hour vodcasts? There are quite a few long series I would far prefer to read than listen.
A more interesting idea would be a browser extension that lets you open a chat window from within YouTube, letting you ask it questions about certain parts of the transcript with full context in the system prompt.
We're at Emergent Mind are working on providing bits of a technical transcript to a model and then asking follow up questions. You can check it out here http://emergentmind.com if curious.
https://chatgpt.com/share/66e9f5ae-8d20-8000-b3a5-7c1ba928b8...
https://chatgpt.com/share/66ea22ad-5d20-8009-a3b0-909c5f500a...
https://chromewebstore.google.com/detail/asktube-ai-youtube-...
I tried pasting the URL of a YouTube video and I get the message "I'm unable to access the video directly, as the tool needed for that is disabled. However, if you'd like, you can summarize the video or let me know how I can assist with it!"
This is what I'm getting: https://chatgpt.com/share/66ea4f36-90b4-8009-8b6c-02bc26cff9...
"I was unable to retrieve the transcript for this video due to its large size."
Here's Lex 8-hour Podcast about Neuralink https://textube.olivares.cl/watch?v=Kbk9BiPhm7o&format=txt
https://chatgpt.com/share/66ea502e-935c-8009-a9f3-5ce9173e57...
https://www.youtube.com/watch?v=fgYIFiWgBl8
It looks like its currently limited to Android phones.
I’ve done this by manually copy/pasting a yt transcript into chatgpt (and later streamlining it into a bash function), and it was quite effective, allowing me to dodge a couple of click bait time wasters. (videos that looked important but really were just fluffing up unimportant nonsense).
I just run yt-dlp to fetch the transcript and shove it in the GPT prompt. (I think also have a few lines to remove the timestamps, although arguably those would be useful to keep.)
My prompt is "{transcript} Please summarize the above in bullet points"
The trick was splitting it up into overlapping chunks so it fits in the context size. (And then summarizing your summary because it ends up too long cause you had so many chunks!)
These days that's not so important, usually you can shove an entire book in! (Unless you're using a local model, which still have small context sizes, work pretty well for summarization.)
Are you sure you're looking at automatic transcripts? YouTube transcripts are bizarrely low quality if they're not provided by the creators (I've actually used my Google Pixel's live transcription to make better captions occasionally).
I just checked a video my girlfriend uploaded a week ago and the auto-transcript was still pretty messy. I've used Whisper for the same task and it's significantly better.
I know people who upload a video on YouTube unlisted just to get transcripts generation for free and then delete the video.
whisper v3 large on release day was around 1x on a 4090
Whisper is only a few tenths of a cent per hour transcribed if transcribing on your gpu though, at about 30x real-time on a 3080 etc. with batching.
do you have a source? more generally is there a community or news source for youtube "api" news like this?
Compare the results:
TexTube: https://chatgpt.com/share/66e9f424-32c4-8009-b761-c8a8d6fbec... VoxScript: https://chatgpt.com/share/66e9f443-31d8-8009-b396-dba11b2f5b...
But if I start a new session and simply paste the link to the video it gives the transcript. I’m not sure an llm is the best solution to getting full transcripts.
I.e. what are the kind of things I can ask and get value from?
They copy paste text transcripts into an Llm and have it generate more text based on its training and prompt data. You can't "chat" with a text document of course.
chatting about a text document
Chatting with a text document implies it has AI or magical abilities.
You wouldn't say you are chatting with your dog if you are talking to your wife about your dog.
It can be useful; it's not hype nonsense.
So rather than watch the video or read the transcript you just ask the one thing you want to know.
Could it take you to the moment in the video that is useful too?
Another solution would be to skip the LLM prompting part altogether and
1. break the transcript into short sections
2. create embeddings from them and remember the timestamps for each
3. embed your query (what are you interested in)
4. calculate the closest embedding in the transcript to your query
5. return the original timestamp
Currently, with the API [1], you can retrieve a JSON with timestamps. The main issue, though, is how to parse the text effectively into meaningful sentences, and then add the timestamps at the beginning of the paragraph. WIP.
[1]: https://textube.olivares.cl/watch?v=9iqn1HhFJ6c&format=JSON
This is an interesting example, it feels different than watching the ~12min video. https://chatgpt.com/share/66e9eaff-248c-8009-9761-d848d97881...
asking questions to transcript at least is ground based on something real (a video)
https://chatgpt.com/share/66eadbad-1d3c-8009-91f0-abe3cf4d36...
Also since in order to find a feature through a literal string you first have to guess it... But language is inherently fuzzier, so literal searches are in this purpose weaker than an interface dealing with the fuzzy aspect of expression.
This is actually what inspired us to create Lectura: https://lectura.xyz/
We’ve added features that promote curiosity and deeper learning, like ELI5 explanations, suggested queries based on transcripts, quizzes to track retention, and more.
If you’re interested in joining us to build out the platform, feel free to reach out at neil at lectura dot xyz