Bypassing YouTube video download throttling
blog.0x7d0.dev
blog.0x7d0.dev
Last time I read discussion about it in yt-dlp repo [1], you can actually bypass it by just adding range=xxx query parameter (not header), and it will return to full speed even if your range is just the whole thing.
And IIRC YouTube have already lifted this restriction.
Edit: find the ref [1] https://github.com/yt-dlp/yt-dlp/issues/6400
Almost all of them have the same protection: some code that triggers only when you open the tools and stops the video by creating a debugger statement you cannot skip and triggering some cpu-heavy code (probably an infinite loop, although I wouldn't discard cryptominers). More importantly this code also clears the network request information, making it more difficult to analyze the traffic sent so far. Note to Firefox devs: enabling "persist logs" should persist the logs. Don't clear them!
None of this is perfect and I never found a video I couldn't eventually download (timing attacks ftw), but I do wish I could find a deeper explanation on how this all works.
If you missed it, not so long ago there was a submission that evaded exactly this. Their solution is so simple yet effective: Recompiling the browser with the debugger keyword renamed. Made me smile.
Our computers are our realms. God giveth and god taketh away.
The "disable breakpoints" button is seemingly performing a one-time imperative action on toggle (maybe "find all the breakpoints + debugger statements in existing loaded code, and patch/unpatch them") rather than being a declarative state-change in the system (i.e. "disable this tab's access to the debugger/breakpoint logic itself").
Not exactly what I'd call simple...
Definitely a clever hack
0: https://chrome.google.com/webstore/detail/anti-anti-debug/mn...
For console.clear I have this
window.console.clear = function clear() {console.log(clear.caller, "Asshole script called console.clear()")}
hoping I get a glimpse of what exactly is calling it.In Firefox, when a console is cleared, it's not just the display that's cleared. The entire console object is reset. This means that any modifications made to the console object, such as redefining console.clear(), are lost when console.clear() is called.
You can try this on this demo [1] website that uses devtools-detector library [2]
Also, the "Open With"[0] browser extension can detect a video/frame's URL when you right click on it, if you want to use that to quickly open it in yt-dlp/mpv.
0: https://addons.mozilla.org/en-US/firefox/addon/open-with/
I've seen this technique too, but I feel that this is a major flaw on the browser's side. It should be impossible to tell if the dev tools is open or not. Surely this can be done right?
All (3) Browser engines are anti consumer spyware. There is no way to convince me otherwise. Why google and apple are not fixing this is obvious but Mozilla just keeps disappointing.
I'm desperately waiting for a Foss alternative.
You didn't truly believe this was about UBI did you? We solved that one ages ago with bank accounts and KYC.
This really has puzzled me. I downloaded a few favourites and watch them on VLC or Infuse on my AppleTV. In the YouTube app I can use the "nerd stats" to confirm I am viewing the exact same video/audio streams...
... it could be my imagination but it seems like YouTube does a really subtle kind of filter that makes the "blocky" compression artifacts smoother. It doesnt ehance edges or anything - my guess is it looks for areas WITHOUT edges where there are sublte shfts of colour, and it makes the blocky artifacts less prominent.
... it's really subtle and I still cant tell if its just my imagination, like my OCD thinking that my downloaded video doesnt look as good and yet, I noticed on YouTube the video feels more vibrant and solid. When I watch my downloaded vid there are these really sublte, but noticable artifacts often in the background , in the shadows and these constant tiny little jitters even on a 1440p video - make the final picture look not as good.
Am I making this up?
Audio wise there is definitely a change as well. YouTube audio is always more or less level for me, while a downloaded video always needs to crank up the volume which is annoying.
I wish players like VLC or Infuse did whatever YouTube does to make videos just more pleasant to look at. I dont think YoUTube changes the colours or does any kind of vibrancy filter though I may be wrong, but it does things to "level" audio so that you have a more consistent experience going from one video/channel to another.
Why don't you just take a screenshot at the exact same timestamp? It's that easy and would have taken you less time than writing this up.
I'm not sure what the defaults currently are for youtube-dl and its forks, but for a long time it defaulted to the best combined stream. However the best distinct audio and video streams are higher quality.
Colour shifts (as in, input != output, not gradients) can come from bad handling of video colour space or monitor profile. Also, shenanigans here can have screenshots looking different from the actual application.
'jitters' might be dropped frames, but then you mention resolution. Since you also mention edges, if you're noticing pixellation in the edges of coloured objects, that would be nearest-neighbour chroma upscaling, which I do remember some player using at some point.
The downloaded video is also 1440p, same audio/video streams as far I can tell. So both Infuse and YouTube will do some scaling to the viewport 1440 > 1080p.
This is an example video:
https://www.youtube.com/watch?v=-VE-tgVOZN8&t=3m30s
So the left side is really dark and that's the area where you'd typically see more compression artifacts right? Due to the algorithm thinking there's no detail there. So in YouTube it feels solid. But in Infuse I notice tiny little jitters there and it just distracts and I'm guessing it's those really subtle "grainy" things that take away from the picture feeling really clean and smooth.
Now when I switch back to YouTube and I really look for it, at same time stamp I can notice some artifacts, but it's just not as noticable... so I'm still wondering what is going on. Does YouTube also do some kinda brightness/contrast filter perhaps similar to audio? Due to playing back on a TV maybe?
Without being able to make screenshots it's really hard to tell since the time it takes to switch between the apps you get flashes or brightness/dark and the eyes are affected by it. All I can tell is in YouTube the picture just feels smoother and cleaner overall.
edit:
Another example gives some hitns perhaps
https://www.youtube.com/watch?v=-VE-tgVOZN8&t=6m32s
So now I am checking out the video on my desktop linux with a 1440p monitor, in YoUTube and in VLC player (Ubuntu, AMD GPU).
So interestingly in the background behind Ally to the left on VLC you can clearly see the banding of blue colours, there are lots of squares which are jittering, like the fuzzy grain on a night cam. It's really distracting.
In YouTube (same desktop, via Google Chrome), there is colour banding in the blue background to the left of Ally, but it is not as noticeable because it's like the squares have been averaged and the edge of the bands is smoother. While you can see some tiny shaking there in the colour bands if you look for it, it is not distracting from the overall image.
Hmm.
UPDATE /SOLVED?
Ok after redownloading 1440p 271 stream (VP9) I can confirm the color banding is the "smooth" one I saw in YouTube.
Something's fishy with YouTube I did download the VP9 codec ~3 weeks ago and since then they have changed the streams and the bitrates are lower. There are these new "6xx" streams while the old 1xx/2xx streams appear to have a lower filesize.
But oddly enough the 400 MB filesize I just downloaded has the smoother nicer picture, whereas the 500 MB file I downloaded weeks ago, has the squares/fuzzy/grainy effect. VLC tells me both are VP9 so hmm.
Interestingly since I have the older download from ~3 weeks ago I just ran mediainfo on it, and the new download from today. Then I just switch tabs in terminal so I can easily see what changes.
Old / New
vp09 / vp09
filesize 488 mb / 391 mb
OVerall bitrate 3274 kb/s / 2620 kb/s
Bits/(Pixel\*Frame) 0.036 / 0.028
I don't see anything else significant.It seems to me in recent weeks YouTube has changed the streams, added new "high bitrate Premium" streams (edit: WHILE lowering the bitrate on the older existing streams like 137!), and perhaps the one I downloaded earlier, despite being a larger filesize, was not encoded correctly?
The new file despite being 20% smaller (~400 mb instead of 500), has the smoother color banding, and doesn't show the ugly jittery grainy artifacts.
I used to think larger filesize is better but I guess Iḿ going to get the Vp9 from now on...
I'll remember that even two+ days after the initial upload, there may still be reencoding. I thought this was only few hours after the creator uploads.
So the smaller filesize also is not necessarily a sign of lower quality. In this case it's because it gets a quick first encoding like zip with fast compression, and then it gets reencoded presumably by something much more CPU intensive.
That said from 500mb to 400mb is a bit dodgy but what do I know maybe VP9 is that good.
When encoding videos there's a "speed" parameter that basically tells the encoder to spend more CPU time compressing the frames. It results in a smaller output size while maintaining the same level of quality but takes much longer to encode. I'm sure that's part of what YT is doing here, an initial quick encode to get the video live then another pass to reduce filesize for long term storage. Good find!
I'm guessing that the perceived higher quality in the smaller file is because the encoder was able to find a more accurate way to represent those frames with the same or fewer number of bits since it has more time to search for the optimal encoding.
.
I've never seen such a long format list in youtube-dl before. Are 6xx new? Apparently,[1] they were introduced together with that 'premium 1080p' this April.
Comparing older 4k videos to your video: this one[2] now has 6xx, and 4xx are gone, and curiously the reported bitrates of all streams have since changed (reencoded?). 137 stayed about the same this time, but 18 dropped from 730k to 493k. For this video[3] 4xx are still available.
.
616 is not the actual premium 1080p as the posters at [4] think, is it? Currently youtube.com chooses 248 when playing [2], but yt-dlp can list and download 614 and 616 without any account cookies.
Rather, 6xx seem to comprise (of) vp9 spanning a medium, high, and sometimes low bitrate in all resolutions. Is yt considering replacing the older formats with these?
I just hope the original 18 and 22 remain for older videos, where any difference in quality also matters the most. When still available, in most cases the H.264 streams with creation_time prior to ~2013 are dramatically clearer than any more recent formats.
[1] https://github.com/yt-dlp/yt-dlp/issues?q=605+604+603+sort%3...
[2] https://github.com/yt-dlp/yt-dlp/issues/1863#issue-106877303...
[3] https://github.com/yt-dlp/yt-dlp/issues/389#issuecomment-103...
[4] https://github.com/yt-dlp/yt-dlp/issues/6770
lets not discuss how often the h264 streams for new videos are higher quality than the vp9
YouTube (and most other streaming sites like Spotify etc) use something called ReplayGain. It's essentially a tag that specifies the calculated average loudness of the video/song/whatever (this number is calculated at upload time).
Upon playback, the official YT client knows to use that tag and adjust its volume level accordingly, but I'd imagine either the tag isn't getting downloaded, or perhaps MKV doesn't support ReplayGain tags natively.
Volume / Normalized 100% / 100% (content loudness -0.2dB)
You got me thinking now. I see there is some kinda ReplayGain postprocessor plugin for yt-dlp, however it's tied to some music downloading. I wish yt-dlp had a builtin option of some sort to process that tag.It's disgraceful that even major movie studios often do such a bad job with audio mixing that I indiscriminately run everything through a filter. This is not a problem with the files I'm using; I find the same thing in the theater. C'est la vie.
In addition to what the other poster said about checking if you're downloading the exact same codec etc as you're watching, you could try playing the downloaded video with the browser and see what it looks like.
In particular, I like the anime4k shader pack, which is ML-based but runs in real time in mpv (and I think VLC as well). While it is tuned for anime (as is obvious from the name), it has decent denoise and deblur which often make YT content more watchable and a restore step that does a really good job with compression artifacts but is a bit too tuned for anime so may not always work, or even make things worse. See https://github.com/bloc97/Anime4K/releases
I'm not [ADVERTISEMENT] sure, because my YouTube viewing [ADVERTISEMENT] experience nowadays is so [ADVERTISEMENT] [ADVERTISEMENT] frequently interrupted with ads that it [ADVERTISEMENT] breaks my focus. #pleaselikeandsubscribeandclickonthenotificationbell
Imagine we could have Internet Premium and never see ads again.
For a while I didn´t consider the "full" Premium because I was always shown the 17 € version - which is for family.
Now suddenly I'm being shown there is a 12 € option which is for single person.
Still for someone like me who doesnt care about music or movies, I'll stick to the Lite version.
- the issue I was experiencing is with a lower quality encoding vp9 stream - from a video that was recently uploaded (explained below)
- though I didn't trust smaller filesize VP9 streams initially, they look in fact noticably better - the picture is smoother, cleaner, the artifacts of compression are less visible. Where AAC can have jittery/glittery distracting dots moving in the background in areas where you have subtle gradients (eg. a plain wall) - Vp9 has none of these, those areas look smoother and cleaner and it gives an overall nicer looking picture without compromising the detail as far I can tell
- Opus audio stream appears to have less of the ReplayGain issue, I'm not sure - but since I downloaded Opus instead of the 140 m4a stream I notice I dont need to adjust the volume compared to viewing same video in YouTube - and since the codec is newer anyway and the filesize is relatively the same or a tad smaller - also it is in 48k not 44k, I am going to download Opus from here on
- a very confusing thing is it appears ; for a recent upload ; you can have an initial VP9 stream of say 500mb which is in fact no better than the AAC and havs the grainy artifacts - and the vp9 stream gets replaced weeks later by one significantly smaller like 400mb vs 500 mb !! and looks way better . whic hsuggst there was a first pass with low quality encoding, replaced by a higher quality encoding later - therefore my assumption that larger filesize is better was wrong
What they care about is you wasting their bandwidth. For an ad-supported video streaming site, bandwidth is normally more expensive than revenue - Google only manages to make it just about work because they have probably the worlds cheapest bandwidth due to being able to bully ISP's into peering with them for free. (they don't let you peer with Google for just Google Search but not youtube).
All these throttling measures are simply trying to reserve most of the bandwidth for real users, not people scraping all the content.
They don't do it yet, probably because they don't see the need quite yet. But i have no doubts that it will happen sooner or later.
Adblocking will have to evolve to a new level to block such things.
I suspect that google doesn't actually lose too many to blockers, as mobile accounts for a large fraction of youtube's traffic (and so far, not that many people actually use a hacked youtube client to view videos).
It's probably cheaper and faster to have a pre-encoded video, cached at the edge.
Anyway, when they start delivering ads in-band, the next step for blockers is to identify that first keyframe in the player by using a pool of shared signatures, right? So then player clients will need adblock plugins which will have a sizeable signature distribution infra and grief for clients.
Then the anti-blocker might begin adding, per-play instead of per-video, a pixel or something to throw off the signatures, massively increasing THEIR video distribution infra. Ad infinitum?
AI controlled adblocker is the end game!
ah yes, AAI (Artificial-AI) AKA I (Intelligence), or "Crowdsourcing" if you're looking to use an older buzzword. I do think there's a few models trained on sponsorblock already, but they're not great.
Many VCRs could do that, and stop/start recording to skip ads, as that was the only way to do it.
Edit: Or make them unskippable?
I think Twitch users tend to watch on their computer mostly? And I think Twitch viewers are more techy so they would be more likely to have ad blockers.
I have no data on any of this. I'm just throwing shit at the wall.
- Introduces a decryption step, which is slow
- Forces software video decoding, which is slow
- Web browsers only support the weakest form of Widevine which is ineffective
It would effectively push a significant portion of their user base off the platform while not being very effective in its goals.
No it doesn't. Heck, forcing software decoding is actually one of the ways to force Widevine down to lower protection levels on general purpose hardware.
> - Web browsers only support the weakest form of Widevine which is ineffective
It's not fullproof but it would certainly make tools to bypass it clearly illegal in most of the world.
There are efficacy reasons for not doing it on a backend level, but Google has required anything that wants YouTube to support Widevine for a very, very long time now.
https://www.legifrance.gouv.fr/codes/article_lc/LEGIARTI0000...
If you're using a PC with a Nvidia GPU, run `nvidia-smi dmon -s u` and start playing a random Youtube video in Chrome. You'll notice how dec% moves from 0% to at least 2%. Pause, and start playing Widevine protected video and notice how dec% stays at 0% because decoding is happening on the CPU.
> It's not fullproof but it would certainly make tools to bypass it clearly illegal in most of the world.
Good luck, copyright infringement is already illegal and yet that hasn't stopped it from being widespread. Tools and techniques to bypass Widevine L3* are widely known and available (yes, even on GitHub).
I was being generous in my previous comment. In reality, deployment of Widevine L3* should be shunned at least as much as Proof-of-work cryptocurrencies. It's completely ineffective in protecting content, it burns unnecessary CPU cycles multiplied by (potentially) billions of users, and significantly degrades user experience.
Even Widevine L1* is ineffective in practice. Techniques to bypass it aren't available to the average Joe, but of course there are groups that will download, decrypt, and re-upload the newest 4K streaming releases to torrent trackers within an hour of them appearing on streaming services.
*edit: Mixed up L3 and L1
It's because Widevine have embedded decoder into its lib and its using CPU instructions but from user perspective it's not a huge change on modern CPUs as most have specialized instructions to handle decoding of H264 etc.
> Widevine L1* is ineffective in practice. Techniques to bypass it aren't available to the average Joe, but of course there are groups that will download, decrypt, and re-upload the newest 4K streaming releases to torrent trackers within an hour of them appearing on streaming services.
There are no "Techniques to bypass it", the only way currently to get L1 streams is to use legit hardware keys from some devices, on which you can exploit secure enclave/extract HW keys.
There are no "instructions to decode H264", there is dedicated hardware acceleration like Intel QSV and AMD VCN, but these gets bypassed just like Nvidia's decoding acceleration from my previous example. All of this is trivially observable, playing back DRM-protected video wastes an obscene amount of resources, relatively speaking.
From user perspective you'll notice stuttering, unusually high CPU usage, dropped frames and more, especially once you try to play multiple videos at once.
> There are no "Techniques to bypass it", the only way currently to get L1 streams is to use legit hardware keys from some devices
That's exactly what I meant. Being pedantic over my choice of words isn't very productive.
For L3 you are just using SIMD/vector instructions compiled for specific platform, so they are specialized CPU instructions (not general use) that help with decoding. And L3 is mostly now 720p and 1080p low bitrate on majority of streaming services that people use, you would need to have VERY old hardware to not be able to use it. I've been watching 720p/1080p h264 videos 15 years ago with only CPU decoding without ANY issues, most of the world did. So that's just not an issue. If we are talking about L1 then you have hardware acceleration so your point is invalid in that case.
> From user perspective you'll notice stuttering, unusually high CPU usage, dropped frames and more, especially once you try to play multiple videos at once.
Yeah because 99% of people are playing multiple widevine videos at once on their 20 years old hardware... come on.
> That's exactly what I meant. Being pedantic over my choice of words isn't very productive.
Im not being pedantic, you are not bypassing a lock in a door with a key, don't you? "Hey honey lets bypass our neighbour door lock using his key so we can enter his house" No one says things like that. If you meant what I meant then you just used wrong words to describe that. Your choice of words have different meaning which isn't very productive.
It's not slow, after you decrypt the AES key, then you are using hardware AES instruction set supported by most CPUs currently.
So like 99.99% of use cases?
> but I promise you that you're not going to have a good time if you try to play back multiple high bitrate videos on slightly older hardware (think HTPC).
That could be the case, not disagreeing here, if you use old hardware, have multiple high bitrate videos and multiple streams at once but this is specialized niche example.
It also paints a very big target on Widevine Level 3's back.
But ultimately it's just a financial equation. What Google are losing from ad blocking isn't quite worth pulling the WV lever yet, but given it has clearly become enough to take softer measures and the pressure they are likely under from music labels I expect a wider rollout will happen in the next few years.
If I were to predict it will probably initially be "any video containing label music or studio clips picked up by content ID", and maybe an opt in tag for other creators at first. They're the ones that are much more useful for monetisation anyway, and you don't lose all your CDN benefits at once.
Using the hardware based version could cause a lot of problems with unsupported devices.
For professional content, there is simply not much demand to develop a user-friendly way to break L3 widevine, as the market is already served by reasonably convenient illegal streaming sites and torrents that allow for higher quality video.
(My use case is not so much in downloading a video, I just want to be able to my Amazon Prime videos on my Linux computer at a better resolution than 720p.)
I used results from this Github search page (especially pywidevine) a while back to make a Hulu downloader in Go which I eventually had to remove: https://github.com/chris124567/hulu
I wasn't really a part of this scene but there appears to be some sort of weird competition among people involved in writing this kind of software so occasionally device keys would leak when they tried to get at each other which was great for me because it meant I didn't have to extract keys from a phone or NVIDIA shield or anything annoying like that.
There are relatively easy ways to download L3 content, but they are not as common because higher resolutions are available on illegal streaming sites and torrents. These sites break hardware Widevine or use other attacks to get the material.
1. Majority of ISPs do not host any cache for Google content
2. Credible ISPs do not have bottlenecks at the transit or peering level
3. Netflix makes use of much more local caching but their model works very differently to Youtubes
4. The concept of "internet backbone" does not really translate to reality. Peering is significantly more mesh-like than that, and transit more diverse.
Source: I have owned multiple ISPs, and still do.
Even if you do it from a browser, it is only accessible from the YouTube website run offline as a SPA.
Even then, downloaded videos can only be played for 29 days before having to reconnect most of the time, with some regions restricting it to 48 hours.
It’s not the raw video file, you can only access it from the app or website, and there’s restrictions on how long you can be offline before the app/website will stop you from watching it.
I tried to trick google by creating an account using VPN in Europe. I even managed to subscribe for Premium. Even with an active Premium subscription Youtube app won't let me use its features such as downloading, background play and picture in picture (you know, basically everything)
Because "you live in the wrong part of the world". Creating a problem, selling the solution is what they do.
Like, I don't even care about downloading, but Google intentionally cripples their mobile website experience by suppressing pip and background play. It would cost zero dollars not to do this but they did.
The fact Apple allows Google to resell their multitasking and PIP features as part of their own subscription is pretty un-Apple.
Ethically, if you don't only think “fuck Google”, I feel like it's reasonable to stop after the first optimization (“pass the real browser test to get regular browser speeds”). There you're not “wasting” any more of YouTube's resources than a browser user with ad-block.
Getting full Gb/s without paying anything feels to me like you're pushing all the ad-blocked users' luck.
But then again, fuck Google I guess?
I think the original browser use case is tuned for the common occurrence of not watching the whole video. But if you intend to watch (and archive) the whole video to begin with then I don't think this eats away Google's bandwidth more. OTOH it probably has more overhead due to the amount of connections.
It's weird from the standpoint of someone who sets out to watch an entire thing, but almost nobody sticks around while watching videos. It seems like the mass of youtube viewerdom are bouncing around like mad, all the time.
Google can't only think about humans, it's got to also think about competing organizations. That complicates things.
It technically is, but if someone makes a programmer's salary and still jumps over turnstiles, I consider it a big red flag over their personality.
In the end you're downloading a lot more data in total, as what ends up happening is you're downloading all sorts of stuff that you never end up getting around to actually watching (data hoarder syndrome). Whereas if you could only watch that stuff during life playback, you'd be downloading much less data in total as there's no bytes wasted amassing a library that never gets viewed.
Well that's an assumption on the intent of downloading. I don't download videos that I don't eventually watch. But typically it's for archival, I just hate my favorite videos going away.
I doubt it.
Companies do not have ethics, only interests. It's only natural I behave similarly.
While that's certainly true (although an simplification - they are just managed by people with low ethics), we can do better. If I behaved like managers of Google, Meta or Microsoft, I'd be ashamed of myself.
No. I shouldn't have to work harder myself to somehow cancel out the evil in the world. Because if we were less cowards and calling them for what they are - evil - they wouldn't hold so much power over us.
Reminds me of that scene from Mr. Inbetween about bullies. Bullies exist because we're told to ignore them and get punished if we retaliate, so they get away with it.
--
But without getting sidetracked. Companies have no emotion, no particular sense of ethics, and least of all, they don't need people to defend them, unless they're called a lawyer.
Not only you should, but you should also support countermeasures that make it harder for evil, so you eventually wouldn't have to work harder.
If there is one reason, do it for your kids, to show them that a better world is possible.
Exactly, if we held public demonstrations outside the offices of book publishers, Sony, RIAA, and similar greedy bastards, and it became the norm to snub our noses at their employees then things would soon change.
For instance, we ought to be demonstrating in the streets over how these bastards are hounding the Internet Archive, but we're not.
If we were, then these companies would quickly change their tune and think twice before launching such lawsuits.
Trouble is we're not out there demonstrating. And it's only a tiny minority of the population who actually care about such things—those of us posting here on HN etc.—who do. We're such a small force we couldn't escape from a wet paper bag on the deck of a sinking ship let alone take on the might of these greedy corporations.
Cory Doctorow has said this many times although he's not been as blunt about it as I am. I've followed this for decades and I reckon it's essentially a lost cause.
Even if we could get politicians to agree to change laws they could only do so around the edges as they've signed international treaties, Berne, WIPO, etc. which prohibit signatories from exiting. Any country that left the treaties would have sanctions placed against it.
These corporations have not only won but they've implemented a system that's irreversible, like a ratchet, every one of their cog-like actions squeezes us consumers further and there's fuck-all we can do about it.
To do otherwise plainly makes you shit as well.
A person should be ethical to another person, because that person can reciprocate. A person should only interact with a corporation in terms of legal frameworks, which are also amoral, because that is the only "moral" framework within which a corporation can act. The fictional personhood of a corporation is just that, a fiction.
(There is somewhat of a sliding scale on this, in that a small corporation formed to protect a fruit stand or something is certainly not Microsoft and shouldn't be treated as a Microsoft.)
three years of covid nonsense -> crickets
video download speed throttle -> rage
Something is wrong here.
(That said, there really was some real nonsense, such as certain dumb countries that penalized people for going on bicycle rides in rural areas by themselves.)
There's no good-for-society reason behind throttling video download speeds.
Google certainly would disagree. Their argument (I assume) would go something like: if people don't pay for content up front and also don't watch the ads then these services can't exist, therefore if these services existing is better for society than them not existing, then it follows that <DRM, etc., fill in the blank> is good for society.
You might not like it but a great deal of our economy is built on that premise, so a lot of people and companies have a stake in holding onto such arguments.
Such an argument is a serious and honest one, and should be responded to with some care. Outright dismissal is not interesting.
> That's because there was a very good reason for that Covid "nonsense" as you put it: reducing infection rates and keeping healthcare systems from collapsing.
That definitely turns out to have been false. The models were vastly wrong. We definitely have differential handling of the pandemic, from Sweden, Africa, and some states in the U.S. for example, and those that went all out did not do better than those that didn't.
Many people warned that this was overblown, but also many people greatly enjoyed exercising authority, and others greatly enjoyed a sense of moral virtue ("saving grandma") that was unjustified.
If your morals go away when they actually matter, you didn't actually have morals, you just had an excuse why you weren't already rich. A person is judged by what they do, explicitly and especially when it actually matters.
Are you meaning towards everyone? Or are you meaning your behaviour back towards those companies?
You wouldn't treat a deranged serial killer with respect and courtesy. Why should poorly-behaved amoral corporations be treated as you would treat normal humans?
Corporations should be treated in accordance with their own behavior. The mom-n-pop shop down the street that uses an LLC for legal purposes and treats you like a valued customer? You should treat them with respect and kindness. The evil megacorp that tries to lobby for shitty laws to screw you over? You should screw them over too. (And alternatively, the big corporation that makes good products for decent prices and doesn't seem to be actively trying to blatantly harm society and twist things to their advantage with legal tricks or lobbying? You should treat them respectfully too.)
I'm using a browser extension to show full size of images on social media when hover and Instagram went crazy about it, warning about unusual activities and threatening to lock my account. I looked into it to see if the extension does any scrapping behind but no, it all seems fine. My assumption is that Meta detects requests happening in wrong order which might indicate data scrapping due to the non-standart client behaviour.
So, measures and countermeasures. I guess it's up to YouTube to implement their counter measures and push the scrappers to implement theirs but overall at some point they should be able to limit downloading to the playback*2 speed because a legitimate consumption wouldn't happen any faster on a legitimate client.
But honestly, downloading shouldn't be restricted. The contents are protected by laws anyways and most people have legitimate reasons for downloading videos. It could be for archiving purposes because the video has some value for you, it could be for analysis or it could be about creating content based on the vide content(like downloading a movie trailer to extract parts for you movie review video). The content in YouTube is not created in vacuum and once published they create the new environment where new videos will be made and this creative freedom is important unless we want the current videos be the last videos ever made.
I've used such an extension and found dragging the curser over your screen might fire like 25 requests. You might simply be rate-limited.
Removed the extension anyway
> Getting full Gb/s
Technically most of the time video is served on the same network(ISP) as youtube has caches with almost all ISPs, IXes in the world. It might be the largest CDN built to date.
Case in point, the authors YT URLs point to Canadian ISP Videotron.
Do you know for example that you cannot take screenshots of netflix in your browser on Mac, and therefore cannot make a meme, which is a fair use that has been taken from you?
ah i remember this one: https://news.ycombinator.com/item?id=32793061
i confess i still dont really understand why they had to make this but i'd love to hear the story behind it
There is an option now to select the audio track and MrBeast uploads dubs in more than a dozen languages.
0: Jimmy's voice in the spanish dub of his channel is the same actor who dubs spiderman.
https://blog.youtube/news-and-events/multi-language-audio-mr...
Especially in an article that already had to unwind and explain Youtube's Javascript code.
This reminds me of some sort of fizzbuzz test. This is not complicated at all. There is no need to use the Range header or run Javascript.
The short script below does not download anything because there is no need. It does not use Range headers, it does not run Javascript and it makes only one TCP connection. With the JSON it fetches, one can simply extract the videoplayback URLs and put them in a locally-hosted HTML page with no Javascript.
#!/bin/sh
# usage: echo videoId | $0 <-- this will indicate len to use
# usage: echo videoId | $0 len | openssl s_client -connect www.youtube.com:443 -ign_eof
# usage: $0 len < videoId-list | openssl s_client -connect www.youtube.com:443 -ign_eof
(
while read x;do
test ${#x} -eq 11||continue
if test $# -ne 1;then len=${#x};x=$(grep -m1 ^\{ $0|sed 's/\$x//'|wc -c);exec echo usage: ${0##*/} $((x+len));fi
cr=$(printf '\r');
sed "/^[a-zA-Z].*: /s/$/$cr/;s/^$/$cr/" << eof
POST /youtubei/v1/player?key=AIzaSyA8eiZmM1FaDVjRy-df2KTyQ_vz_yYM39w HTTP/1.1
Host: www.youtube.com
Content-Type: application/json
Content-Length: $1
Connection: keep-alive
{"context": {"client": {"clientName": "IOS", "clientVersion": "17.33.2" }}, "videoId": "$x", "params": "CgIQBg==", "playbackContext": {"contentPlaybackContext": {"html5Preference": "HTML5_PREF_WANTS"}}, "contentCheckOk": true, "racyCheckOk": true}
eof
done
printf '\r\n'
printf 'GET /robots.txt HTTP/1.0\r\nHost: www.youtube.com\r\nConnection: close\r\n\r\n';
)
For processing the JSON I wrote custom utilities in C that (a) extract videoIds and other useful strings, (b) generate HTTP similar to above, and (c) filter the returned JSON into CSV, SQL or HTML. For me, these run faster than Python and jq and are easier to edit. Using these utilities I can also do full searches that return hundreds to thousands of results and I can easily exclude all "suggested" or "recommended" videos.CSV output
1666520150,23 Oct 2022 10:15:50 UTC,22,aqz-KE-bpKQ,"Big Buck Bunny 60fps 4K - Official Blender Foundation Short Film",00:10:35,635,UCSMOQeBJ2RAnuFungnQOxLg,19211597,"Blender"
SQL output
INSERT INTO t1(ts,utc,itag,vid,title,dur,len,cid,views,author) VALUES(1666520150,'23 Oct 2022 10:15:50 UTC',22,'aqz-KE-bpKQ','Big Buck Bunny 60fps 4K - Official Blender Foundation Short Film','00:10:35',635,'UCSMOQeBJ2RAnuFungnQOxLg',19211597,'Blender') ON CONFLICT(vid) DO UPDATE SET views=excluded.views;
HTML output
Looks just like CSV except vid is a hyperlink
echo -n '{"videoId":"aqz-KE-bpKQ","context":{"client":{"clientName":"WEB","clientVersion":"2.20230810.05.00"}}}' |
http post 'https://www.youtube.com/youtubei/v1/player' |
jq -r '.streamingData.adaptiveFormats[0].url'
which is very similar to what you do, but runs into an issue of throttling to ~70Kbps. Is the difference just the "key" parameter? Do you get no throttling?it's not the key the author is using, it's the post-data.
Moreover, to get throttled videoplayback URLs with the "WEB" key and client info like the author is using, one does not need to make POST requests to /youtubei/v1/player. There are throttled videoplayback URLs in the HTML of the /watch?v= page. For example,
curl -A "" -40s https://www.youtube.com/watch?v=aqz-KE-bpKQ|grep -o https://rr[^\"]*|sed -n 's/\\u0026/\&/g;/itag=22/p'
It's ironic how the author is claiming this is complicated. That's his own doing. curl 'https://...googlevideo.com...' --output video.mp4
for me the download is throttled to "768k", i assume thats in bits per second and not bytes which is very low: the random video i tried would take 8 minutes.on the other hand,
yt-dlp videoIdHere
does its processing then downloads the whole thing in about 5 seconds.Does that curl command run much faster for you? Or do you do something else?
Use this post-data and should get same speed as yt-dlp.
{"context": {"client": {"clientName": "ANDROID", "clientVersion": "17.31.35", "androidSdkVersion": 30 }}, "videoId": "$x", "params": "CgIQBg==", "playbackContext": {"contentPlaybackContext": {"html5Preference": "HTML5_PREF_WANTS"}}, "contentCheckOk": true, "racyCheckOk": true}
I do not use curl, except in HN examples. I generally do not download from YouTube. I use the URLs in the JSON to watch the video.However, acquiring this key requires decompiling the mobile application, monitoring requests through a proxy, or relying on values discovered by others. It's not necessarily straightforward.
I do agree that the code is simpler this way.
I also find it interesting that, by default, yt-dlp calls the YouTube API three times, initially as an Android client, then as an iOS client, and finally as a Web client. Depending on the video and certain other parameters, YouTube provides different formats to different clients.
This is again not true. The key is in the HTML of every /watch?v= YouTube page. It's a public key; it's not hidden in any way.
Further, it's possible, up until today at least, to use the "WEB" key with clientName "ANDROID" of "IOS" and receive unthrottled URLs. The key in the shell script is in fact the WEB key. The key for IOS is different.
curl -40s https://www.youtube.com/watch?v=aqz-KE-bpKQ \
|grep -o \"INNERTUBE_API_KEY...[^\"]*\"Downloading the same file used in the official UI and applying/reverse engineering the function is not exactly rocket science.
And the other approach is currently in progress - soon logged users won't be able to view videos without watching(or at least displaying/downloading) adds, so the logical next step is to nerf anonymous(as is without google account) viewing. No matter how I look YT has exact problems (and solutions) as all file locker sites. The only difference is that YT is not at mercy of Ad companies, it is the Ad company(at least Google is). So they might try some more aggressive measures, that normally would get a site banned from publishing ads.
Edit: I did some searching. In addition to that, the only relevant news I could find were about Google testing locking the 4K (2160p) resolution behind the Premium doors, which they ceased doing.
But it seems I'm mistaken [1].
[1] https://www.theverge.com/2023/2/23/23612647/youtube-1080p-pr...
Sounds pretty bad for ad-revenue....
I hate to say this because I'd be affected but that's the math I would expect them to use. The only counter argument I can think of, why they might care to keep freeloaders there, is the network effect
Huge amount of users of youtube are literally babies, they don't know how to read or write so they don't sign in.
I'm starting to also suspect the trend of YTs algorithm insisting on recommending videos you've already watched is also a decision from this huge baby user base.
1. Solve the challenge the same way the browser would. By actually running the JS code.
2. Segmented downloading. Pretty sure they allow that on purpose so starting or resuming a video feels snappy.
Also for switching to a different bitrate stream seamlessly.
That's pretty impossible with open-source software. Google engineers aren't clueless; they'll know where to find this information just as well as anyone here.
It's too bad there isn't a legal way of preventing Google engineers from reading some source code.