Youtube-dl
rg3.github.io
rg3.github.io
I remember starting the project around 2006. Back then, I had a dial-up connection and it wasn't easy for me to watch a video I liked a second time. It took ages. There were Greasemonkey scripts for Firefox that weren't working when I tried them, so I decided to start a new project in Python, using the standard urllib2. I made it command line because I thought it was a better approach for batch downloads and I had no experience writing GUI applications (and I still don't have much).
The first version was a pretty simple script that read the webpages and extracted the video URL from them. No objects or functions, just the straight work. I adapted the code for a few other websites and started adding some more features, giving birth to metacafe-dl and other projects.
The raise in popularity came in 2008, when Joe Barr (RIP) wrote an article about it for Linux.com.[1] It suddenly became much more popular and people started to request more features and support for many more sites.
So in 2008 the program was rewritten from scratch with support multiple video sites in mind, using a simple design (with some defects that I regret, but hey it works anyway!) that more or less survives until now. Naturally, I didn't change the name of the program. It would lose the bit of popularity it had. I should have named it something else from the start, but I didn't expect it to be so popular. One of these days we're going to be sued for trademark infringement.
In 2011 I stepped down as the maintainer due to lack of time, and the project is since then maintained by the amazing youtube-dl team which I always take an opportunity to thank for their great work.[2] The way I did this is simply by giving push access to my repository in Github. It's the best thing I did for the project bar none. Philipp Hagemeister[3] has been the head of the maintainers since then, but the second contributor, for example, was Filippo Valsorda[4], of Heartbleed tester[5] fame and now working for Cloudflare.
[1] http://archive09.linux.com/articles/114161 [2] http://rg3.name/201408141628.html [3] https://github.com/phihag [4] https://github.com/filosottile [5] https://filippo.io/Heartbleed/
I vastly prefer offline media players to browser-based tools, for a number of reasons: better controls and playback, richer features, uniform features (I don't have to learn each individual site's idiosyncracies), the ability to queue up a set of media from numerous sources and play them back without clobbering one another, and more.
Hugely useful tool, and I've been impressed as hell as well by its update frequency.
And lift a mug to old Warthog. I miss Joe as well.
Why don't browsers provide some way to play local video files, for example by typing "file:///c:/my_video.flv" into the address bar. After all, the browser certainly includes the ability to play the video being downloaded off the web.
If you try "file:///c:/my_video.flv" with Firefox, it opens a dialog box offering to pass the video file to whatever external media players you have installed.
In what seems inconsistent to me, "file:///c:/my_notes.txt" and "file:///c:/my_pic.jpg" will be rendered correctly by Firefox -- it won't offer to open an external text editor or photo viewer. Why is video different?
Also, what always impressed me is the incredible amount of random contributions from the community. Ever since we introduced a super-simple plugin system [0], support for the most disparate video sites poured in as PR. (>800 PR!!) Also, given how ytdl is structured, the most simple plugin gets you 90% of the tool power for that video sit. Big results with minimum effort.
Finally, to answer the question about the updates in some siblings, there is no active effort against us most of the time (VEVO videos being the notable exception) but supporting such a number of sites mainly by scraping means that breaking changes happen really really often.
[0] https://filippo.io/add-support-for-a-new-video-site-to-youtu...
Regardless, youtube-dl and VLC complement each other quite well.
Do you think DRM support in browsers and major tube sites will soon prevent tools like youtube-dl from functioning?
And I must say I'm impressed by its ease of use (basically zero installation effort), and also by the frequent updates.
(I wonder why those frequent updates are necessary, though. Are you under the impression that google is actively working against tools which attempt to download material from youtube?)
As @fillipo said above, there is little if any pushback from video sites. Most of the time, they update their interface (we've gotten better in anticipating minor changes) and something breaks. The recent string of YouTube breaks (for some videos, mostly music videos - general video is unaffected) is caused by the complexity of their new player system, which forces us to behave more and more like a full-fledged webbrowser. But I think we usually manage to get out a fix and a new release within a couple of hours, so after a small youtube-dl -U (Caveats do apply[0]) you should be all set again. Sorry!
Anyway, if you didn't write this tool (and update it) -- I'd have to do it myself. And I'd rather not do anything myself ;-)
As the catalyst and original dev for this tool, Thank you!
I setup a makefile to let me just go make and then eventually vlc pops up with stuff to watch every so often. Its quite nice.
# this script uses sh, sed, awk, tr and some http client
# here, some http client = tnftp
# awk and tr are optional
# wrapper for tnftp to accept urls from stdin
ftp1(){
while read a;do
ftp ${@--4vdo-} "$a"
done;}
# uniq
awk1(){ awk '!($0 in a){a[$0];print}' ;}
# some url decoding
f1(){
sed '
s,%3D,=,g;
s,%3A,:,g;
s,%2F,/,g;
s,%3F,?,g;
s/^M
//g;
# ^ thats Ctrl-V then Ctrl-M in vi
'
}
# remove redundant itags
f0(){
sed -e '
s/&itag=5//;t1
s/&itag=1[78]//;t1
s/&itag=22//;t1
s/&itag=3[4-8]//;t1
s/&itag=4[3-6]//;t1
s/&itag=1[346][0-9]//;t1
' -e :1
}
# separate urls
f2(){
sed '
s,http,\
&,g'
}
# remove unneeded lines
f3(){
sed '
#/^http%3A%2F.*c.youtube.com/!d;
/^http%3A%2F.*googlevideo.com/!d;
/crossdomain.xml/d;
s/%25/%/g;
s,sig=,\&signature=,;
s,\\u0026,\&,g;
/&author=.*/d;
'
}
# separate cgi arguments for debugging
f4(){
sed '
s,%26,\
,g;
s,&,\
,g;
'
}
# remove more unneeded lines
f5(){
sed '
/./!d;
/quality=/d;
/type=/d;
/fallback_host=/d;
/url=/d;
/^http:/!s/^/\&/
/^[^h].*:/d;
/^http:.*doubleclick.net/d;
/itag.*,/d;
'
}
# print urls
f6(){
sed 's/^http:/\
&/' | tr -d '\012' \
|sed '
s/http:/\
&/g;
'
}
f8(){
sed 's/https:/http:/'
}
FTPUSERAGENT="like OSX"
case $# in
0)
echo|$0 -h
;;
[12345])
case $1 in
-h|--h)
echo "url=http[s]://www.youtube.com/watch?v=..........."
echo usage1: echo url\|$0 -F \(get itag-no\'s\)
echo usage2: echo url\|$0 -g \(get download urls\)
echo usage3: echo url\|$0 -fitag-no -4o video-file
echo N.B. no space permitted after -f
;;
-F)
$0 -g \
|tr '&' '\012' \
|sed '
/,/d;
/itag=[0-9]/!d;
s/itag=//;
/^17$/s/$/ 3GP/;
/^36$/s/$/ 3GP/;
/^[56]$/s/$/ FLV/;
/^3[45]$/s/$/ FLV/;
/^18$/s/$/ MP4/;
/^22$/s/$/ MP4/;
/^3[78]$/s/$/ MP4/;
/^8[2-5]$/s/$/ MP4/;
s/.*?//;
'|awk1
;;
-g)
while read a;do
n=1
while [ $n -le 10 ];do
echo $a|f8|ftp1||
echo $a|f8|ftp1 &&
break
n=$((n+1))
done \
|f2|f3|f1|f0|f4|f5|f6|f1|sed '/itag='"$2"'/!d'
done
;;
-f*)
while read a;do
n=1
while [ $n -le 10 ];do
echo $a|$0 -g ${1#-f} |ftp1 $2 $3 $4 $5 ||
echo $a|$0 -g ${1#-f} |ftp1 $2 $3 $4 $5 &&
break
n=$((n+1))
done
done
;;
esac
esac
There are separate scripts for extracting www.youtube.com/watch?v=........... urls from web pages to feed to this script.Personally I have no need for "VEVO" videos. Nor do I ever encounter VEVO youtube urls posted to websites, like HN. I wonder why?
As for maintainability, I beg to differ. The raison d'etre for this script arose out of frustration that early YouTube download solutions, e.g. gawk scripts, clive, etc., kept breaking whenever something at YouTube changed. I got tired of waiting for these programs to be fixed, if that ever happened.
I can fix this 164 line script faster if YouTube changes something than waiting for a third party to fix something they developed that is far more complex. Moreover, it does not rely on Python. Is there something wrong with DIY?
I see someone posted a link in this thread to another 208 line script, yget, that uses sed and awk. This further demonstrates the relative simplicity of downloading YouTube videos.
Go to youtube.com, you may need to scroll but most unlikely, bam! "VEVO" video with x-million views: it's a music video promotion brand.
Actually it's not globally promoted so outside of Western Europe and USA I'd guess you don't get VEVO vids so much?
According to https://www.youtube.com/watch?v=5zs1ClgqhLw their 100th most viewed video has 200 million views. Top 10 are all above 600 million.
They're quite a big brand.
An alternative to goofing around on the youtube.com web site, scrolling constantly and getting hit with advertising and endless lists of "related" videos is to search and retrieve youtube urls from the command line via gdata.youtube.com.
[1]: https://github.com/rg3/youtube-dl/tree/master/youtube_dl/ext...
As I recall, it was originally written by one person (Ricardo Garcia) in 2008 and worked only on YouTube using (by later standards) relatively simple heuristics to find the URL to extract the video. But it's catalyzed an explosion of interest in every aspect of the problem: tracking changes to the HTML of the video sites, adding support for more video sites, figuring out indirection and parsing through multiple pages and HTML objects, making the tool much more multiplatform and easier to install and update...
It's attracted hundreds of contributors (many of them motivated by a personal desire to be able to use the tool on a different site, or to fix a bug that was preventing them from downloading video in a particular rare case) and maintained an incredibly rapid pace of development.
This kind of project that requires a lot of fairly laborious work to create support for many different information sources is a particularly good candidate for an open source project.
It's the opposite of what you've stated.
Authors, composers and publishers needed protection against cheap printing presses that would just print anything that was popular and flog it in the marketplaces.
The limitations I mention are the difficulties and cost of the printing itself.
What authors wanted was to restrict who can print their work -- but it's not true that authors "needed protection" because printing presses started appearing.
That makes it sound like authors were paid for the work until those "cheap printer" pirates appeared. But on the contrary it was the invention of the printing presses themselves that gave authors an industry in the first place -- for millenia authors just wrote for free.
Yes; a good read is "The Surprising History of Copyright and The Promise of a Post-Copyright World" [1] which I think is from Karl Fogel, the author of the (Free, Libre, CC-BY-SA) book "Producing Open Source Software" [2]
[1] http://questioncopyright.org/promise [2] http://producingoss.com
"it's not unlike an anti-capitalist punk rocker STEALING her clothes at H&M".
That said:
First, I fail to see the contradiction from being an "anti-copyright freedom fighter" and "downloading stuff from YouTube".
Someone somehow convinced you than anti-copyright people only like copyleft works? The very idea of being anti-copyright is wanting to abolish all copyright.
Second, what's witht the "anti-copyright freedom fighter" strawman? As if someone needs to be that to want to download stuff off of YouTube?
Might as well have written "like it or not, the South's engine is slavery" in 1850...
The one place I can see where this breaks down is in advertisements, but I consider that to fall into the incidental results. (Although Youtube-dl does have a --include-ads option)
You have opportunity, that's not the same as a right. The content supplier is under no obligation to provide content to you, ergo no "right to see" that content.
That said, personal time-shifting and format-shifting should IMO be a normally allowed part of the copyright deal.
Porn has the need (the mainstream providers generally delete porn) and the sheer resources to do it.
I've just done the config file setting you mentioned.
[0] http://bugs.python.org/issue1739468 [1] https://github.com/rg3/youtube-dl/blob/640743233389714dda8a3...
youtube-dl: youtube_dl/*.py youtube_dl/*/*.py
zip --quiet youtube-dl youtube_dl/*.py youtube_dl/*/*.py
zip --quiet --junk-paths youtube-dl youtube_dl/__main__.py
echo '#!$(PYTHON)' > youtube-dl
cat youtube-dl.zip >> youtube-dl
rm youtube-dl.zip
chmod a+x youtube-dl
Might be less confusing if you append '.zip' in the first two commands: zip --quiet youtube-dl.zip youtube_dl/*.py youtube_dl/*/*.py
zip --quiet --junk-paths youtube-dl.zip youtube_dl/__main__.py
When you echo the shebang overwriting the file, I was thrown off. I'm thinking,
"Why did you just zip all those contents into the file to just throw them out?"
Then I see the `cat` line, and it makes sense that the `zip` command appends
the .zip to the end of the file. ytplay() { youtube-dl "$1" -o - | vlc - }
As a side benefit it of course also allows you to instantly watch stuff from all the other sites YT-DL supports :) ytplay() { youtube-dl "$1" -o - | vlc -; }
but works great! ytplay() { youtube-dl "$1" -o - | vlc - ; }
or: ytplay() { youtube-dl "$1" -o - | vlc - #<enter here>
> } #Where "> " is bash prompting for more/end of definitoin
In a file (eg: .bashrc), I'd personally prefer: ytplay() {
youtube-dl "${1}" -o - | vlc -
}
Note that there's very little difference between "$1" and "${1}" in practice, I tend to prefer it for consistency with recommended[1] practice of using ${NAME} rather than $NAME. (And to differentiate something like "${1}${2}" vs "${12}", as you might if $1 was a name, and $2 an extension, or $1 and url-scheme and $2 a host-name (http://hostname -> "${1}${2}" 1="http://" 2="hostname").[1] http://stackoverflow.com/questions/8748831/bash-why-do-we-ne...
youtube-dl --get-url "$1" | xargs mplayerIf not, please share me a YT link.
$ youtube-dl -citw ytuser:LastWeekTonight
I downloaded a channel with 121 videos, 4.4 gigs, took 26 minutes, so 2.8MB/s average. Curious if the Youtube people will shrug it off and free the beer or rate limit or more aggressively combat this.
Also, to get the total number of supported sites:
$ youtube-dl --extractor-descriptions|wc -l
466 (wow)
As this can run on anything with Python, I guess that includes Android[0], iOS[1], Windows Phone[2], heck even Blackberry[3]??
[0] https://python-for-android.readthedocs.org/en/latest/
[1] https://code.msdn.microsoft.com/windowsapps/using-python-on-...
[3] http://forums.crackberry.com/blackberry-z10-f254/blackberry-...
Thanks pmoriarty for submitting this. Awesome and I'm just getting started poking around with it. Makes me really want to learn Python, seems that's what all the fun stuff[4] is coded in.
[4] http://motoma.io/pyloris/ :)
[0] https://github.com/rg3/youtube-dl/blob/master/README.md#do-i...
$ youtube-dl -citw ytuser:UCxIJaCMEptJjxmmQgGFsnCgIt can also convert a video to an mp3:
youtube-dl --extract-audio --audio-format mp3 https://www.youtube.com/watch?v=OKbtC223e30
> youtube-dl -F OKbtC223e30
139 m4a audio only DASH audio 49k , audio@ 48k (22050Hz), 525.88KiB (worst)
171 webm audio only DASH audio 129k , audio@128k (44100Hz), 1.27MiB
140 m4a audio only DASH audio 129k , audio@128k (44100Hz), 1.36MiB
172 webm audio only DASH audio 176k , audio@256k (44100Hz), 1.78MiB
141 m4a audio only DASH audio 255k , audio@256k (44100Hz), 2.71MiB
Just use youtube-dl -f 141 to direct download just the audio from YouTube. I always use this to download songs. --audio-quality QUALITY
"best", "aac", "vorbis", "mp3", "m4a", "opus", or "wav"; "best" by defaultffmpeg.exe -i "Keith Wiley - The Fermi Paradox, Self-Replicating Probes, Interstellar Transport Bandwidth-AUk6ZlePtQA.m4a" -c:a copy 2.m4a
to change the format in the file header to M4A.
(Edit) Answering my own question, apparently yes!
youtube-dl -x --audio-format mp3 OKbtC223e30
youtube-dl -x --audio-format mp3 -fCtvurGDD8By the way, the GitHub issue tracker (https://yt-dl.org/bug ) is usually a better place to report issues. But just for youtube-dl reaching #1 on HN, I'll make an exception.
youtube-dl -x --audio-format mp3 -- -fCtvurGDD8 youtube-dl -x --audio-format mp3 -- -fCtvurGDD8I have messed around with one sketchy youtube downloader or another for years until I found this.
Oh, and it's on brew!
It'll also download an entire playlist, and add sequential numbers at the beginning (with the -A option).
It's very nice though for legitimate developers.
Seriously though, awesome project, used it for a while.
[1] http://caca.zoy.org/wiki/libcaca [2] https://www.gnu.org/licenses/license-list.html#WTFPL
I mean I understand that all the separate steps are stuff we've seen is easily possible nowadays, but putting it together in a single UI makes it (roughly) 3000x as useful!
I want to try it, anyone had luck compiling it for Linux?
EDIT/update: Well I gave it a try, grabbed QtCreator, loaded the project, not much luck. Some issues with the "phonon" library, it seems. I'm not very good with getting C++ stuff to work when it gives build errors. I did spend about half an hour fiddling and googling error messages, but now it's time to give up, sorry :)
I'm writing this update to let you know that one of the errors I did manage to fix, is that Windows has case-insensitive paths/filenames, while Linux does not. Apparently the path for the phonon library is lowercase, so you should `#include <phonon>` lowercase. I'll try to leave a Github issue about this.
That didn't help much (complaints about the State enum in soundfix.h) which I tried to fix by also putting `#include <phonon>` in the soundfix.h. I'm not sure if that was right at all, but it did seem to fix that particular problem. As a result I was greeted with a whole bunch of other (I think unrelated?) errors about some types not being strictly compatible or something. That is where I check out until I know more about C++, decided it had been long enough, and just writing you a little message to let you know how it went.
However, it made me install and try QtCreator, something that I was meaning to do anyway. So that's a win :)
And a php web app that uses youtube-dl as a a backend: https://github.com/Rudloff/alltube/
http://www.verticalforest.com/youtube5-extension/
Helped to get rid of Flash, although I have Chrome in my programs folder for sites that are still flash only.
function youtube-dl {
exe=$HOME/bin/youtube-dl
link=https://rg3.github.io/youtube-dl/download.html
url=$(curl -s $link |grep ">sig<" |head -1|sed -e 's/href="/|/g' -e 's/">/|/g'|cut -d"|" -f2)
fetch=N
if [ -s "$exe" ] ; then
ts=$(date "+%s")
yts=$(stat -c "%Y" $exe)
[ $(( ($ts-$yts)/(60*60) )) -gt 24 ] && fetch=Y
else
fetch=Y
fi
if [ "$fetch" = "Y" ] ; then
url=$(curl -s $url |grep ">sig<" |head -1|sed -e 's/href="/|/g' -e 's/">/|/g'|cut -d"|" -f2)
echo "Fetching [$url] and deploying to [$exe]"
curl -s $url -o $exe
chmod a+x $exe
fi
[ -z "$@" ] && $exe --help || $exe $@
}At some point I decided to write something similar in Ruby ( https://github.com/rb2k/viddl-rb ) and I'm kind of ashamed of how broken things are from time to time.
Video hosting sites don't have APIs and reverse engineering the sources for the videos is like shooting at a moving target.
So kudos for leading that project :)
youtube-dl seems to solve this: https://github.com/rg3/youtube-dl/issues/2165
Example:
quvi -vm --format $format "$url" --exec 'wget %u -O %t.%e'
instead of $format put any format that the video supports. To query them, use: quvi --query-formats "$url"
So it's going to be something like: quvi -vm --format fmt43_360p "$url" --exec 'wget %u -O %t.%e'
And to extract audio from the result you can use: avconv -i something.webm -vn -acodec copy something.ogg
Youtube however is switching away from fixed video files to separate streams to be used with MSE. You can note that higher resolution video is not available the old way. So downloading that won't be so straightforward.Same question for clive/cclive...
By the way, FFmpeg has support for the 0.4 branch of libquvi, so if you built it with --enable-libquvi you can ffmpeg -i http://youtube... (assuming libquvi and its scripts still work)
Good question, I didn't keep track of its development. Your best option would be asking developers if they are planning to work on it further or not.
> By the way, FFmpeg has support for the 0.4 branch of libquvi
Same as mpv I think. You can play Youtube videos with it directly:
mpv "$url"
Which is kind of fun, since you can do tons of things that aren't available in the browser player - looping, playing only portions of the video and all other things which mpv can do. mpv --ytdl "$url"[1] https://gitorious.org/get-flash-videos-plugins/pages/Hulu
Take a look at the "new" section!
I've downloaded a whole youtube channel, downloaded videos as mp3's and what not.
I have couple of aliases set in my .zshrc too :)
mp3dl() { youtube-dl --extract-audio --audio-format mp3 $1 }
root@haseebr7 ~ mp3dl <youtube_video_url>
and bam i have the mp3 download. I no longer have to visit shitty ad infested websites to do these kind of things.
Getting youtube-dl to run has been a pain for some reason, for example I always seem not to have the right version of Python available. Yget doesn't depend on such volatile tools.
youtube-dl --username pimlottc :ytwatchlater
If you get a problem, please file a bug report at https://yt-dl.org/bug . Thanks!