Youtube-dl is possibly dead
old.reddit.com
old.reddit.com
That of course doesn't mean companies won't sue, or pay off a developer, or make threats, etc..
I'm quite interested about where it sits wrt 'time shifting' which is allowed in the UK. I'm allowed to record "tv" to watch later, AFAIK; so I can watch on the train, or when I'm away from the "tv" receiving equipment (eg aerial) as long as I only watch it once and don't show other people (!). So, can I still download a YouTube video to watch in the car where I _personally_ don't have the ability to watch? YouTube is TV. Or would it only count for a live broadcast show on YouTube?
This is my personal opinion and is not legal advice.
This is gold. What if you mirror it on two TV screens, but one has a slight latency, because HDMI->SCART something something connector?
The TV licensing auditors conferred amongst themselves. "If he sees everything twice, he needs two licenses" one argued. "Hold on" said another, "he reported seeing six shows, he needs six licenses". "We showed him three, he saw six, he needs nine licenses" said another. The second one's eyes narrowed. "The law doesn't allow him to see more than one, he needs a fine". "If he's creating new feeds without permissions, that's distribution and copyright infringement, he could be facing a custodial sentence for this". They turned to Yossarian who decided it had gone far enough. "I see everything once!" he cried. "Destruction of evidence now, is it?".
It's Content-22: if you don't want to watch it, there's no money to be made from you so you can - better that you watch and talk about it than not. But if you do want to watch it, then you can't, not without paying. The content is made for people who want it, but those are the very people who have to be kept from it. The people who want to come in, can't, and the people who want to share it, mustn't.
Yossarian thought about Content-22. He picked up a book. "Put me through to cancellations, I have been meaning to do more reading lately". The TV licensing auditors hesitated and glanced at each other. "One license per person, covers all screens, after all you can only watch one at a time" said one. "We can give you a discount if you extend your license for another year?" said another. "And send you an upgraded set-top box for being a loyal customer".
My understanding is that this is the critical 'gotcha' that necessitates a license for using software in EU. Copyright was supposed to govern copying, not use. Of course the scope of copyright has been only expanded since then.
You walk into a store, you buy a disc, it's yours. Legally obtained copy. No terms.
After that you'd be free to use it without having any terms be imposed on you by the copyright holder, because using isn't copying, and if you're not doing something that is the copyright holder's exclusive right, you don't need their permission to do it.
If we wanted to follow this line of thinking, maybe you should also need a license to listen to music because today's digital media players tend to make temporary copies of the audio in RAM and buffers..
I think this technicality makes a big difference.
Of course I'm sure they could figure out other ways to give copyright holders the power to impose their terms on users, but the way it's currently done in EU feels like a technical 'gotcha' to me.
"Article 4 of Directive 91/250 lays down the exclusive rights of the copyright holder, rights of a preventative nature, in its computer program. The first of those rights is the right of reproduction, which is defined in particularly broad terms because it covers not only any form of reproduction, whether permanent or temporary, but also acts of reproduction necessary to use a program. Unlike other categories of works, in any case those which are distributed on their own medium, a computer program always requires a reproduction, if only a temporary one, in the computer’s memory in order for that program to be used. The rightholder’s exclusive rights therefore constitute, as far as computer programs are concerned, greater intrusion into the private sphere of the user than in the case of other categories of protected subject matter, because those rights require de facto the authorisation of the rightholder even simply to use the program.
..
Thus, the copyright holder’s exclusive rights in a computer program cover not only traditional acts of exploitation of the work under copyright, but also the enjoyment of that work in the user’s private sphere."
https://curia.europa.eu/juris/document/document_print.jsf;js...
"Subject to the provisions of Articles 5 and 6, the exclusive rights of the rightholder within the meaning of Article 2, shall include the right to do or to authorize: (a) the permanent or temporary reproduction of a computer program by any means and in any form, in part or in whole. Insofar as loading, displaying, running, transmision or storage of the computer program necessitate such reproduction, such acts shall be subject to authorization by the rightholder;"
To be clear the "gotcha"s as referred to in the comments above aren't "bits of the written law one might find surprising or overreaching" they are attempts at circumventing the meaning of the law by taking it absurdly literally (as in my stop sign example in my sibling comment) and hoping to use that narrow interpretation as legal reasoning instead of paying attention to the actual intent of said law.
As I said the law as written doesn't singularly hinge on the process of loading a program into RAM to allow the requiring license to use software anyways, the software can require those additional terms as part of the installation onto the computer's storage which is clearly a copying process. This is also why the "walk to the store, buy a disc, it's yours. Legally obtained copy. No terms." example is silly - even if we accepted that was legally yours with no usage terms what you have is no terms usage (but not modification or redistribution) of the installer which is going to require you to agree to terms to perform the copy (and probably enter a license key in any semi-modern time). Similarly with the point on music, what makes you think it isn't licensed to a copyright holder? Because it doesn't prompt you when you hit play? CDs give unlimited private playback and backup rights, no public distribution, and no public performance. Modification has certain allowances (known as fair use) just like software but generally is up to the rightsholder. Streaming gives a temporary license for much the same. Digital radio gives you rights to listen live. There is no slippery slope of digital music gaining copyright control through some reasoning from software copyright, it's all already there.
The point I'm making is that lawmakers themselves used a similar absurd literality ("actually your computer makes a copy in ram to run a program") in writing this law.
> As I said the law as written doesn't singularly hinge on the process of loading a program into RAM to allow the requiring license to use software anyways, the software can require those additional terms as part of the installation onto the computer's storage which is clearly a copying process.
Software doesn't inherently require installation, but if it does, it doesn't have to be permanent. I can make a temporary installation in RAM.
> Similarly with the point on music, what makes you think it isn't licensed to a copyright holder? Because it doesn't prompt you when you hit play? CDs give unlimited private playback and backup rights
The copyright law does not make personal playback of music an exclusive right of the rightsholder. It's not the CD giving me any rights, it's that the copyright law never took these rights away from me in the first place. That is why you don't need a license to listen to music that you own a legally obtained copy of. They never applied the absurd technicality of temporary copies to music.
You're welcome to think any parts of the law are absurd (and given the size you probably will think many parts are) but perceived absurdity does not make a gotcha nor does it mean any other perceived absurdity not explicitly in the law now must also be valid or invalid accordingly. If the law had vaguely said "copies of computer programs are protected" and made no provisions towards the intent of protecting the object code in any form elsewhere then later a an argument in court said "and loading it into memory counts as a copy" to try to expand the right past the intent that would be a case of a gotcha.
> Software doesn't inherently require installation, but if it does, it doesn't have to be permanent. I can make a temporary installation in RAM.
How does the installation being done to a temporary location or not change whether the installer has to produce a copy of the software from itself onto the machine? The action the installer does is independent the medium the bytes are stored.
> To be clear the "gotcha"s as referred to in the comments above aren't "bits of the written law one might find surprising or overreaching" they are attempts at circumventing the meaning of the law by taking it absurdly literally (as in my stop sign example in my sibling comment) and hoping to use that narrow interpretation as legal reasoning instead of paying attention to the actual intent of said law. The point I'm making is that lawmakers themselves used a similar absurd literality ("actually your computer makes a copy in ram to run a program") in writing this law.
> As I said the law as written doesn't singularly hinge on the process of loading a program into RAM to allow the requiring license to use software anyways, the software can require those additional terms as part of the installation onto the computer's storage which is clearly a copying process.
Software doesn't inherently require installation, but if it does, it doesn't have to be permanent. I can make a temporary installation in RAM.
> The copyright law does not make personal playback of music an exclusive right of the rightsholder. It's not the CD giving me any rights, it's that the copyright law never took these rights away from me in the first place. That is why you don't need a license to listen to music that you own a legally obtained copy of. They never applied the absurd technicality of temporary copies to music.
In the case of CDs copyright law provides the other restrictions explictly listed but the CD (typically) choses to give you unlimited playback ability. Copyright isn't the only law that dictates what you're allowed to do with the content privately, newer (2000s) laws around DRM for instance provide more of the other rights restrictions of digital content.
Redditors: Yep, this is the end.
There are open pull requests and/or bugs for many websites that aren't being approved. This is rather unusual for youtube-dl, it normally had a release ever two weeks or so.
Using a non-public API is not at all the same as scraping, which refers to parsing a rendered HTML page for the content you want.
Both have this maintenance problem, but one's not a fancy word for the other.
After all, what is a scrapeable HTML page if not a grotesquely convoluted undocumented API with an unstable output format?
The gross inefficiency and low data-to-layout ratio are the key things being expressed through connotations of the word "scrape". To scrape is to extract a small amount of something from a much larger substrate.
To call every query a scrape is to diminish the specificity and utility of the term.
{
"id": 3422,
"title": "My essay about cheese",
"published": "13th August 2021 at 3:45pm",
"abstract": "<p>In which I write about cheese!</p>"
}
And I write code against that which includes stripping the HTML tags from "abstract" and converting the date format in "published" into in ISO datetime... am I writing scraping code?I would argue that I am, even though it started out as a JSON wrapper.
"To call every query a scrape is to diminish the specificity and utility of the term."
Absolutely disagree with you there. I interpret the term "scraping" as "writing code that gathers data from a source that has not deliberately published that data in a usable format". Gathering data from any kind of API fits that criteria for me, since most APIs only give you a subset of the data at a time.
I think the reason I care so much about this is that I coined the term "git scraping" to cover a variant of scraping that uses Git repositories to store the data and track changes over time - and git scraping applies equally to data sourced from APIs as it does to data sourced from HTML pages. https://simonwillison.net/2020/Oct/9/git-scraping/
The term was coined to differentiate how difficult it is to extract data from a format that was patently not intended to efficiently spread raw data to other machines. If that meaning erodes, and it's just yet another way to say an API query, it will be a great loss for the precision of our terminology.
Most APIs are not designed to give you all of the data at once - they exist to serve other purposes, usually involving returning a small subset of the data to power a user-facing feature.
If someone asks me "where did you get those Olympic medal results?" and I say "I scraped them" I think that's accurate vocabulary whether I parsed HTML or gathered them from hundreds of undocumented API calls.
If I had downloaded a neat CSV file from the Olympics website with all of the data I needed in one go I wouldn't feel comfortable calling it scraping.
Re-reading your comment, I think what I'm describing here does actually fit with your "how difficult it is to extract data from a format that was patently not intended to efficiently spread raw data to other machines" definition - except I'm including APIs that return only a subset of the data as part of those inefficiencies in obtaining the raw data.
Scraping is the act of extricating data from the layout and markup metadata meant to make it pretty for humans.
APIs generally don't include any of that, your HTML-in-a-JSON-object example notwithstanding.
I'd have no objection to calling it scraping when you strip those <P> tags, but aggregating the results of several API queries is bog-standard textbook API usage, which we use the term scraping to differentiate from.
That said, I had a look around and the definitions I could find tended to support my interpretation:
https://en.wikipedia.org/wiki/Web_scraping - "Newer forms of web scraping involve monitoring data feeds from web servers. For example, JSON is commonly used as a transport storage mechanism between the client and the web server."
https://towardsdatascience.com/web-scraping-basics-82f8b5acd... - "There are 2 different approaches for web scraping depending on how does website structure their contents." (HTML scraping and API access)
https://realpython.com/beautiful-soup-web-scraper-python/ - "Web scraping is the process of gathering information from the Internet. Even copying and pasting the lyrics of your favorite song is a form of web scraping! However, the words “web scraping” usually refer to a process that involves automation"
The more formal dictionaries (Merriam Webster and suchlike) don't seem to have formed an opinion on this one yet!
What is the upside of using this word in such an oddly vague, expansive way? What happens to those of us who need to convey the original specific meaning we coined it for in the first place?
I call that parsing, which is one of the steps in scraping that may or not be necessary depending on the data source.
I need a term that means "using automation to gather data from the web, when that data has not been published in a way that is suitable for my purposes". Scraping works great for that!
It's not that there's repo inactivity, it seems to be that this is an extremely active repo which saw everything grind to a halt when the admins went dark. That's quite a bit different then just "inactivity".
That's kinda overwhelming though ... imagine that if the maintainer pops up somewhere, suddenly 100 motivated people may chime in "hey please review this important pull request that's been sitting over here for a while".
There are some kinds of open source projects that are prone to this ... some are really not so bad to maintain if you have the right kind of discipline, because they converge on a stable set of functionality and platform compatibility evolves slowly, but some just naturally have endless room for variations and special cases, and as users increase, PRs increase linearly (instead of sub-linearly as you'd hope). I'm thinking in particular of https://github.com/oauth2-proxy/oauth2-proxy (of which I contributed to an older fork)
When a webpage changes layout, youtube-dl needs updated as well.
We're talking mostly about a list of site definitions more than we are core development.
https://github.com/ytdl-org/youtube-dl/issues/23860
And there is also an entire fork that fixes the support for just a single provider, NicoNico, because the maintainers ignored its issues.
https://github.com/animelover1984/youtube-dl
A quote from its README:
All code in this project is licensed solely with the condition that any portion of it is not permitted to be used in the main youtube-dl fork, either directly or indirectly. It is also not permitted to be used in any project that contains contributions from either remitamine or dstftw.
The two users mentioned are or were previously major contributors to youtube-dl.
It seems that youtube-dl was already a dysfunctionally managed project at the time of the lawsuit and happened to ride out on the good PR for a couple of months, before returning to stagnation once again.
To me it sounds like a plugin system would have prevented centralization and the need for forks, but would have made distribution harder for average users.
I maintain a GCC code coverage tool on GitHub, and since GCC doesn't change very often and the feature set of the tool is fairly complete, I sometimes go 6+ months without commits. Usually I don't touch it unless someone opens an issue.
That being said, the project probably could use some reorganization. It requires a lot of community contributions to keep all the extractors maintained so long turnaround times for reviews isn't ideal.
https://old.reddit.com/r/DataHoarder/comments/p9riey/youtube...
Sounds more like maybe the person is sick or something.
They don’t even bother removing spam.
It could be fatigue. Many hobbit projects end this way: after initial excitement wears off (could be a few weeks, a few months or even a few years), people just keep going with the boring routine out of habit for years and years; then the routine is abruptly interrupted by some event, and people realize they might as well just stop. COVID killed a few hobby projects I was involved in or followed this way, for instance.
Gotta destroy the ring somehow
Other, non NSA/Microsoft forks should continue development where they left off.
And since no one else seems to be capable of linking outside of the NSA/Microsoft walled garden to such a fork, might as well link mine:
youtubedr (https://github.com/kkdai/youtube, written in Golang) is much faster, to the point `mpv --profile=pseudo-gui $(youtubedr url -q ... URL)` is about as fast as playing a video in YouTube's own interface. Though I assume it doesn't support non-YouTube services unlike youtube-dl, and I don't know if it can be extended to add them.
Why is youtube-dl slow? Is it due to the Python interpreter startup/overhead, or due to code flaws? Can it be sped up, or precompiled to not depend on the complexity of the Python interpreter?
Youtube-dlc and youtube-dlp
It would be nice if someone submitted patch and gets accepted, but don't see the big need.
So far I have seen the same concept on three different contributors, therefore I will agree with another comment that most likely developers are on a summer break.
https://github.com/yt-dlp/yt-dlp
Any other recommendations?
> I say to you that the VCR is to the American film producer and the American public as the Boston strangler is to the woman home alone.