Archiving repositories
github.com
github.com
The only way to discover if a project has a more active fork elsewhere is to go to the network tab and scroll around in the graph... Ideally there would be a way to add a banner at the top of dead repositories with a text like "This project is inactive, but there is a more active fork <here>", as otherwise most people visiting the project's page would have no idea such a fork exists.
I've successfully revived an abandoned project (with an owner who ignored all my attempts to contact them), but only because the primary resource for the project was a wiki on another website, so it involved changing the links there - but this isn't the case for projects whose GitHub page is their main website.
Potential for abuse by hostile forks and ill-wishers would be a big problem and I can't think of any real mitigations.
As is often hotly debated in HN comments, measures of repo inactivity/health are not universally accepted. I can see thresholds in any of these metrics potentially problematic if applied universally.
I do sympathise with the pain-point you're outlining, however!
Anyway, the above is one possible concept for such a feature that I think would be pretty resistant to any kind of abuse. But I'm not necessarily arguing in favor to this specific concept. In fact, given how rare such events would be, the most expedient "implementation" of this feature might simply be a dedicated email address and instructions on how to write up such a request for manual review, along with a stated policy of the minimum period of inactivity before a repo is eligible for such treatment. But my overall point is that with abandoned projects, I think you can slow down the timeline for transfer of maintainership to the point that abuse of the feature becomes essentially impossible.
It would be a good idea to do as you propose: if the maintainer doesn't click a button for X months (and the button is prominent when you're logged in, impossible to overlook), make another fork a primary fork.
Bitbucket has (for local branches) a simple behind/ahead visualization [1] that IMO would be very nice to have on forks!
On top of that, they could integrate something like [2] as I pointed in another comment
[1] https://blog.bitbucket.org/files/2013/10/branch-details.png
It might be useful to promote the "Network" graph a little more prominently once a project is archived.
It would be really great if GitHub made something like that a built-in.
On my side, I've been a Cordova developer for a while and I experienced many dead repos (I think most Cordova devs go native at one point).
Fork reconciliation is an interesting issue but not sure how to implement it to prevent abuse. I once got an email from a guy who created a fork, sent an email to all the owners of the other forks who were ahead, to come up with a community fork, I think that's the best we can have for now.
The main problem is how to signalize to other devs that you're interested in maintaining the fork and accepting PRs. One way could be to just comment on open issues/PRs on the dead repo to say "come to my fork", but when the repo is archived, seems it won't be possible anymore. So it would be actually harder now to create a community fork.
I’m actually dealing with this right now with a project[0] that was forked from an unmaintained repo that wouldn’t accept any pull requests.
Really simple, and it actually works surprisingly well.
I still wish there were a direct way to disable all pull requests, though, to set appropriate expectations. (People are still more than welcome to fork the repo and do what they want, there's a free software license on it. But a pull request is asking me to maintain it, and I don't intend to maintain most of the repos on my GitHub.)
If anybody has a reasonable explanation, I'd be very interested to hear it.
Because it's not just a matter of configuration/input-munging, the way getting to "good search" for text is. The existing solutions for search at scale (ElasticSearch, Solr) just don't work well for a code-document corpus the way they work for a text-document corpus. Trying to trick them into doing so will only get you half-way there, and then you're stuck in molasses if you try to improve from there.
You really have to build "code search" from the ground up, doing things like running language-specific parsers over each codebase and then building tries out of linearized AST node token sequences.
Bitbucket took a different approach: by boosting the definitions matching your search term, the result you want is likely to rank much higher (usually #1) in the search results. Bitbucket's algorithm boosts definitions for a wide range of type categories including classes, functions, enums, structs, and interfaces. Bitbucket prioritized building a code aware search scoped to team and user accounts over a global search functionality. This way, we hope to quickly give users the relevant results they want instead of the hassle of checking out a repo locally and searching using an IDE.
https://blog.bitbucket.org/2017/05/02/introducing-code-aware...
I'd like to be able to hide some of our legacy repositories from code search, because when I search across all of the code that belongs to my organization I don't want to see code from repositories that are no longer maintained. I frequently use code search across organization to see if e.g. a specific library is being used, and archived repositories are not relevant to that use-case at all.
My ideal solution would be the ability to default-to-not-including-archived-repos, but with an option for "run this search against archived repos as well".
I'd actually want to disable PRs without making the repo read only.
2. PRs are basically enhanced issues, so it lets you see the discussion on old PRs. (It would make no sense to make the Issues tab disappear on archive, right?)
3. If a project gets archived with outstanding open PRs, and you want to resuscitate the project with a fork of your own, you can also "rescue" the outstanding PRs by creating new PRs that come from the same branches as the original ones.
Edit - Tried actually doing it and here is the message:
> You will still be charged for this repository. This will not change your billing plan. If you want to downgrade, you can do so in your Billing Settings.
ssh user@rsync.net "git clone git://github.com/LabAdvComp/UDR.git github/udr"
Any repos/projects that are important to me get mirrored to my personal rsync.net account every time I use them.
Sometimes repos disappear, or get taken down ...
The new feature "archives" in the sense that it adjusts the repo to be read-only, so that folks can view it on GitHub without being able to PR or file issues. Like historical preservation.
Your example is "archive" in the sense of "have a backup of", which is something more folks should definitely be doing, but isn't a replacement for GitHub's new feature.
I think the difference is not in the meaning of "archiving" but in the meaning of "repository". Github repository is far wider concept than plain git repo:
> Archiving a repository makes it read-only to everyone (including repository owners). This includes editing the repository, issues, pull requests, labels, milestones, projects, wiki, releases, commits, tags, branches, reactions and comments.