https://github.com/ytdl-org/youtube-dl/graphs/commit-activit...
By itself, using the ANDROID API instead of the WEB API^2 does nothing to affect download speeds. I can block yt-dlp's POST request indicating the client and API name and this has no effect on download speed.
A website that forces users to run Javascript in order to get faster download speeds. This is not a new idea.
1. https://www.youtube.com/s/player/{player_version}/player_ias...
2. From the YouTube video page JSON:
ANDROID API KEY
AIzaSyA8eiZmM1FaDVjRy-df2KTyQ_vz_yYM39w
WEB API KEY
AIzaSyAO_FJ2SlqU8Q4STEHLGCilw_Y9_11qcW8
One of the reasons yt-dl and yt-dlp are so slow is because they do (too) many do other things besides running the Javascript, modifying n and sig, and spitting out a fast download URL. Before they can run the JS, the video page needs to be retrieved, but waiting for Python to start up and do this is slow. A YouTube video page can be retrieved much faster using netcat, outside of Python. YouTube video can be downloaded much faster using an HTTP client like tnftp, directly, outside of Python. YouTube video can be converted much faster using ffmpeg, directly, outside of Python. And so on. These programs start instantaniously when compared to the slow start up time of Python. The Python startup latency is unbearable.
At the very least yt-dl and yt-dlp should accept an already downloaded YouTube video page as input instead of forcing the user to use Python (or Python calling another program) to download the page. The recalculation of "n" and "sig" do not have to occur within seconds of retrieving the video page. There is no need to require the user to use Python for downloading webpages. The values in the YouTube video page are good for a substantial amount of time.
Both scripts have an option to output the new n and sig values in a download URL or JSON containing the download URLs, so the user can use an HTTP client directly, outside of Python. But the user still has to use Python to download each video page. Using an HTTP client directly would be faster.
Using yt-dl and yt-dlp just to output optimal download URLs feels like overkill.
Absolutely nobody thinks optimizing the meta/API processing of yt-dlp & co is worth it. This is exactly why we have high level programming languages that make all of this much easier, instead of trying to write HTML and JS parsing in plain C. Keep in mind these tools support dozens or hundreds or websites, not just YouTube.
If you think rewriting yt-dlp in C is worth it, go right ahead, but you're not going to make it significantly faster; you're just going to make maintenance a much bigger pain for yourself. Pick the right tool for the job. Python is absolutely (one of) the right tools for this job.
(For the record: I use C and Python on a daily basis, and will use whatever language is appropriate for each situation.)
This opinion assumes it is a single job. I see multiple jobs. The number "options" provided by yt-dl(p) gives us a clue.
There is nothing wrong with preferring to use larger, more complicated, "multi-purpose" utilities. There will always be plenty to choose from.
However the idea of using smaller, less complicated, single purpose utilities is not "nonsense". It makes sense in many cases and some users may prefer it.
The statements I make about speed are from day-to-day experience not conjecture.
Plus I can request multiple video pages over a single TCP connection with netcat.
For example, in a single TCP connection, with a list of 30 videoIds, I can get initial sig and n values for 413 videoIds. No wait time for netcat to decompress or startup. Using netcat is quite fast. Then I have utilties written in C to extract URLs from stdin. As such, all I need is a utility to update the sig and n values in the download URLs to make them fast ones instead of throttled.
How long would it take to get sig and n values for 30 videoIds let alone 413 with yt-dl or yt-dlp over a single TCP connection. The startup time for yt-dl/yt-dlp plus the time waiting for the downloading of each video page makes it much, much slower. These are "do-everything" scripts that are as a result quite inflexible.
yt-dl and yt-dlp use Python to run the Javascript functions that modify the initial "n" and "sig" values to make non-throttled download URLs, but a faster language could be used to interpret and run the snippet of Javascript. For example, V8 is written in C++, not Python.
One of the reasons yt-dl and yt-dlp are so slow, IMO, is because they do (too) many other things besides running the Javascript, modifying n and sig, and spitting out a fast download URL. Before they can run the JS, the video page needs to be retrieved. I have self-created utilities for downloading webpages wth TCP clients that are sgnificantly faster and more flexible than yt-dl/yt-dlp. I use these small dedicated programs every day. I am used to the speed. Waiting for yt-dl/yt-dlp to start up in order to download web pages is annoying. The latency is unbearable.
At the very least yt-dl and yt-dlp should accept an already downloaded YouTube video page as input instead of forcing the user to use Python (or Python calling another program) to download the page. The recalculation of "n" and "sig" do not have to occur within seconds of retrieving the video page. There is no need to require the user to use Python for downloading webpages. The values in the YouTube video page are good for a substantial amount of time.
Both yt-dl and yt-dlp have an option to output the new n and sig values in a download URL or as JSON containing the download URLs, so the user can use an HTTP client directly to download video, wthout need to launch yt-dl/yt-dlp. But the user still has to use yt-dl/yt-dlp to download each video page. Using a TCP client directly would be faster.
Using yt-dl and yt-dlp just to output optimal download URLs feels like overkill. It would be nice to have a program that just focuses on running the inecessary JS in base.js in order to output optimal download URLs. Then the user can use whatever programs she wants, directly, for downloading HTML/JSON, extracting URLs, downloading video files, converting video files, etc. Downloading video from most websites is generally easy. I never need a program like yt-dl/yt-dlp. It is only YouTube that plays games with users trying to get them to enable Javascript and be tracked. One need only look at the size of the extractors in yt-dl/yt-dlp as evidence. The extractor for YouTube is 3x the size of the next largest, and over 10-20x the size of most of the others. I want a small utility that just focuses on YouTube. A simpler solution instead of a massive project.
--external-downloader aria2c
--external-downloader-args "--continue --max-concurrent-downloads=3 --max-connection-per-server=3 --split 3 --min-split-size 1M"
(possibly in your config file)Also, neither --format best nor --format bestvideo chooses the best encoding in all cases; they use bitrate as a heuristic for quality, and a less efficient codec can have higher bitrate but worse quality, resolution, or framerate. The workaround for this is specifying --format with an enumeration of every combination of codec, resolution, and framerate in preferred order, which goes like this:
--format "(bestvideo[vcodec^=av01][height>=4320][fps>30]/bestvideo[vcodec^=vp9.2][height>=4320] ...
Here's a full example (hmm... they're using it with yt-dlp, which I thought had fixed this?):https://github.com/TheFrenchGhosty/TheFrenchGhostys-Ultimate...
I think there's a bit of variation in the exact order among the config files found online. If you're goals are archival, consider also retrieving metadata, thumbnail, and subtitles in all languages; I also have in my config the options:
--verbose
--download-archive ./ytdl-archive.txt
--cookies ./ytdl-cookies.txt
--merge-output-format mkv
--add-metadata --all-subs --embed-subs
--write-info-json --write-thumbnail
--no-overwrites --continue
--force-ipv4
(the only remaining workaround for age-restricted videos is to give it cookies extracted from a browser with a real Google account logged in)Very misleading phrasing: you would download in 10 hours a ~10 hours long video, which (of course) could have been downloaded in a fraction of the time.
The throttling has the user download at a speed similar to that required to viewing the video.
format 18 - 640x360 AVC1, V+A : ~50kB/s
format 22 - 1280x720 AVC1, V+A : ~60kB/s
format 137 - 1920x1080 AVC1, V : ~128kB/s
format 400 - 2560x1440 av01, V : ~450kB/s
format 401 - 3840x2160 av01, V : ~850kB/s
Are you sure you are not reporting mega/bits/? When you mention the «player reported data rate», are you sure that is not the "connection speed"? 10MB/s means downloading 36GB/h (one CD per minute, dozens gigabytes per hour)...Typical video bitrates for bog standard 1080p30 will be in the 1MB/s range, so the throttling is around 20x slower than real time.
function ytp() { youtube-dl --get-id "$1" | xargs -I '{}' -P 200 youtube-dl -i --embed-thumbnail --add-metadata -f 'bestaudio[ext=m4a]' -o '%(title)s.%(ext)s' 'https://youtube.com/watch?v={}'; }I use it to download videos into a Plex library.
https://github.com/firedm/ shows no public repos and of course https://github.com/firedm/FireDM gets a 404
https://dereferer.me/?https%3A//www.jwz.org/hacks/youtubedow...
Why?
Just copy the link to not send any referrer information.
It not only completely prevents stuff like this, it profoundly increases your privacy on the web by preventing sites from tracking which domain you came from. There is no good reason any site needs to know that. I am surprised that Mozilla hasn't simply made this the default setting for all users.
I was under the impression it was, I doubt I'm the only one, so thanks for drawing attention to it.
https://old.reddit.com/r/WaybackMachine/comments/kzvzxl/fail...
> it happened to me and I figured out it started when I disabled website.referrersEnabled
So if you completely disable sending the referer header, it breaks. This would probably also happen if you set `network.http.sendRefererHeader` to 0 or 1.
But that's not what I suggested! I only suggested disabling sending the header externally, to other sites, when the host domain (*.example.org) changes. In fact, someone in the thread you linked says that doing this instead of disabling it completely fixed the issue for them:
> ok I left network.http.sendRefererHeader on 2 (had it on 1), and changing network.http.referer.XOriginPolicy to 2 it works
Thanks! Just checked and it is at zero for the default. Reviewed https://wiki.mozilla.org/Security/Referrer and learned something new.
> Connection reset by peer
in a script that I have that downloads podcasts and immediately transcodes to low bitrate opus using ffmpeg.
It's a real bummer...