https://fatcat.wiki/coverage/search?q=is_oa%3Atrue+year%3A%3...
and are working to improve crawling. There is a "save paper now" feature, as well as an API for bots. Organizations like DOAJ, ISSN, DOI registrars (Crossref, Datacite, others) are crucial for this. In the broader ecosystem, we hope this can complement existing efforts that partner with large publishers (like LOCKSS, Portico, JSTOR) and institutional repositories. A natural niche for us is web-native (HTML) content, which we have crawled a lot of but are just getting started to index. For example, publications like d-lib, first monday, and distill.pub.
If folks want to help, it would be great to have a "youtube-dl for open access papers". There is a lot of content on large platforms and publishers which have anti-crawling measures (even for gold OA and hybrid content!), as well as a long tail of small publishers that don't use simple/common mechanisms like OAI-PMH and the `citation_pdf_url` HTML meta tag to identify fulltext content. The OAI-PMH ecosystem sadly is not very complete or helpful for the use case of mirroring.