Redditors: Yep, this is the end.
Redditors: Yep, this is the end.
It's not that there's repo inactivity, it seems to be that this is an extremely active repo which saw everything grind to a halt when the admins went dark. That's quite a bit different then just "inactivity".
That's kinda overwhelming though ... imagine that if the maintainer pops up somewhere, suddenly 100 motivated people may chime in "hey please review this important pull request that's been sitting over here for a while".
There are some kinds of open source projects that are prone to this ... some are really not so bad to maintain if you have the right kind of discipline, because they converge on a stable set of functionality and platform compatibility evolves slowly, but some just naturally have endless room for variations and special cases, and as users increase, PRs increase linearly (instead of sub-linearly as you'd hope). I'm thinking in particular of https://github.com/oauth2-proxy/oauth2-proxy (of which I contributed to an older fork)
When a webpage changes layout, youtube-dl needs updated as well.
We're talking mostly about a list of site definitions more than we are core development.
https://github.com/ytdl-org/youtube-dl/issues/23860
And there is also an entire fork that fixes the support for just a single provider, NicoNico, because the maintainers ignored its issues.
https://github.com/animelover1984/youtube-dl
A quote from its README:
All code in this project is licensed solely with the condition that any portion of it is not permitted to be used in the main youtube-dl fork, either directly or indirectly. It is also not permitted to be used in any project that contains contributions from either remitamine or dstftw.
The two users mentioned are or were previously major contributors to youtube-dl.
It seems that youtube-dl was already a dysfunctionally managed project at the time of the lawsuit and happened to ride out on the good PR for a couple of months, before returning to stagnation once again.
To me it sounds like a plugin system would have prevented centralization and the need for forks, but would have made distribution harder for average users.
There are open pull requests and/or bugs for many websites that aren't being approved. This is rather unusual for youtube-dl, it normally had a release ever two weeks or so.
Using a non-public API is not at all the same as scraping, which refers to parsing a rendered HTML page for the content you want.
Both have this maintenance problem, but one's not a fancy word for the other.
After all, what is a scrapeable HTML page if not a grotesquely convoluted undocumented API with an unstable output format?
The gross inefficiency and low data-to-layout ratio are the key things being expressed through connotations of the word "scrape". To scrape is to extract a small amount of something from a much larger substrate.
To call every query a scrape is to diminish the specificity and utility of the term.
{
"id": 3422,
"title": "My essay about cheese",
"published": "13th August 2021 at 3:45pm",
"abstract": "<p>In which I write about cheese!</p>"
}
And I write code against that which includes stripping the HTML tags from "abstract" and converting the date format in "published" into in ISO datetime... am I writing scraping code?I would argue that I am, even though it started out as a JSON wrapper.
"To call every query a scrape is to diminish the specificity and utility of the term."
Absolutely disagree with you there. I interpret the term "scraping" as "writing code that gathers data from a source that has not deliberately published that data in a usable format". Gathering data from any kind of API fits that criteria for me, since most APIs only give you a subset of the data at a time.
I think the reason I care so much about this is that I coined the term "git scraping" to cover a variant of scraping that uses Git repositories to store the data and track changes over time - and git scraping applies equally to data sourced from APIs as it does to data sourced from HTML pages. https://simonwillison.net/2020/Oct/9/git-scraping/
The term was coined to differentiate how difficult it is to extract data from a format that was patently not intended to efficiently spread raw data to other machines. If that meaning erodes, and it's just yet another way to say an API query, it will be a great loss for the precision of our terminology.
Most APIs are not designed to give you all of the data at once - they exist to serve other purposes, usually involving returning a small subset of the data to power a user-facing feature.
If someone asks me "where did you get those Olympic medal results?" and I say "I scraped them" I think that's accurate vocabulary whether I parsed HTML or gathered them from hundreds of undocumented API calls.
If I had downloaded a neat CSV file from the Olympics website with all of the data I needed in one go I wouldn't feel comfortable calling it scraping.
Re-reading your comment, I think what I'm describing here does actually fit with your "how difficult it is to extract data from a format that was patently not intended to efficiently spread raw data to other machines" definition - except I'm including APIs that return only a subset of the data as part of those inefficiencies in obtaining the raw data.
Scraping is the act of extricating data from the layout and markup metadata meant to make it pretty for humans.
APIs generally don't include any of that, your HTML-in-a-JSON-object example notwithstanding.
I'd have no objection to calling it scraping when you strip those <P> tags, but aggregating the results of several API queries is bog-standard textbook API usage, which we use the term scraping to differentiate from.
That said, I had a look around and the definitions I could find tended to support my interpretation:
https://en.wikipedia.org/wiki/Web_scraping - "Newer forms of web scraping involve monitoring data feeds from web servers. For example, JSON is commonly used as a transport storage mechanism between the client and the web server."
https://towardsdatascience.com/web-scraping-basics-82f8b5acd... - "There are 2 different approaches for web scraping depending on how does website structure their contents." (HTML scraping and API access)
https://realpython.com/beautiful-soup-web-scraper-python/ - "Web scraping is the process of gathering information from the Internet. Even copying and pasting the lyrics of your favorite song is a form of web scraping! However, the words “web scraping” usually refer to a process that involves automation"
The more formal dictionaries (Merriam Webster and suchlike) don't seem to have formed an opinion on this one yet!
What is the upside of using this word in such an oddly vague, expansive way? What happens to those of us who need to convey the original specific meaning we coined it for in the first place?
I call that parsing, which is one of the steps in scraping that may or not be necessary depending on the data source.
I need a term that means "using automation to gather data from the web, when that data has not been published in a way that is suitable for my purposes". Scraping works great for that!
That being said, the project probably could use some reorganization. It requires a lot of community contributions to keep all the extractors maintained so long turnaround times for reviews isn't ideal.
https://old.reddit.com/r/DataHoarder/comments/p9riey/youtube...
Sounds more like maybe the person is sick or something.
I maintain a GCC code coverage tool on GitHub, and since GCC doesn't change very often and the feature set of the tool is fairly complete, I sometimes go 6+ months without commits. Usually I don't touch it unless someone opens an issue.
They don’t even bother removing spam.