PyPy has moved to Git, GitHub
pypy.org
pypy.org
It's kind of sad that this is true.
I'm guilty myself, I contribute to projects on GitHub more often than on any other platform.
And when I search for open source projects the first page I use is GitHub.
That said I do feel some joy when I see a project on Gitlab and am happy to contribute there, eg FDroid.
They had a sizable lead and completely bungled it.
While SF was crapping where they eat, GitHub built a lot of trust and goodwill with a lot of people.
I would consider SF a viable Github alternative, but the bad reputation caused by a temporary owner just seems to stick forever.
That ancient feeling UI doesn’t win them a lot of forgiveness, if the only notable change has been “we took away the malware” and the site continues to remain stagnant, SourceForge will continue to feel very inferior to more modern alternatives.
Some people don't even know the difference between Git and GitHub...
I have also found this to be the case, even with engineers that have years of experience. It's both impressive and awful.
Maybe Copilot (made possible by the huge non-commercial codebase on Github) being somewhat unfair advantage to other commercial alternatives is a bit troublesome, yeah. But otherwise I just don't see why Github being a de-facto standard is bad. In fact, I am somewhat annoyed when a really popular project doesn't have a Github repository (mostly because it makes filing an issue, or even reading existing issues much more difficult in most cases). So I'm actually glad to hear that some big projects feel pressed to migrate to github. What's even a problem with that, apart, maybe, for github actions, that honestly suck?
(Maybe I should add: I am a git hater, and do think that mercurial is just unquestionably better, but this battle is lost long time ago, so I don't suppose it's the topic of this discussion.)
Edit: They have C-layer emulation, but I don't know its limitations or current status, but you can use those libraries [1][2]
[1] https://www.pypy.org/posts/2018/09/inside-cpyext-why-emulati...
Given the way that pypy is implemented, I think the name is quite clever really.
> one can argue it is bad name
I suppose one can, but it's a python interpreter written in python, so I think it's pretty good.
I would have thought I'd need this much more, but I have not. In plain git I'll just `git log` and grep for the commit in case I want to make sure a commit is available in a certain branch.
What practical use case am I missing out on when these work-in-progress draft commits are lost? I can’t see one.
"“When we actually examined the behavior and looked for new attack vectors, we discovered that if you download a malicious package — just download it — it will automatically run on your computer,” he told SC Media in an interview from Israel. “So we tried to understand why, because for us the word download doesn’t necessarily mean that the code will automatically run.”
But for PyPi, it does. The commands required for both processes run a script, called pip, executes another file called setup.py, that is designed to provide a data structure for the package manager to understand how to handle the package. That script and process is also composed of Python code that runs automatically, meaning an attacker can insert and execute that malicious code on the device of anyone who downloads it." https://www.scmagazine.com/analysis/a-third-of-pypi-software...
By all means, I prefer Git branches.
> The difference between git branches and named branches is not that important in a repo with 10 branches (no matter how big). But in the case of PyPy, we have at the moment 1840 branches. Most are closed by now, of course. But we would really like to retain (both now and in the future) the ability to look at a commit from the past, and know in which branch it was made. Please make sure you understand the difference between the Git and the Mercurial branches to realize that this is not always possible with Git— we looked hard, and there is no built-in way to get this workflow.
> Still not convinced? Consider this git repo with three commits: commit #2 with parent #1 and head of git branch “A”; commit #3 with also parent #1 but head of git branch “B”. When commit #1 was made, was it in the branch “A” or “B”? (It could also be yet another branch whose head was also moved forward, or even completely deleted.)
In this post they say that "Github notes solves much of point (1): the difficulty of discovering provenance of commits, although not entirely"
There is either base branch A whose current head is commit #2 / branch B with head of commit #3.
OR
Commit #1 is branch “default” commit #2 is branch “A” with parent as commit #1 and commit is branch “B” with parent also as commit #1
Consider your same example with forking instead of branching, how would the issue be resolved?
The question isn't how many branches, it's what branch the commit was on at the point in time it was created. That's not up to interpretation. It's information that was not recorded.
> Consider your same example with forking instead of branching, how would the issue be resolved?
Forked repositories don't have IDs, don't generally keep track of each other, and there's no way to even count them. So that's not solvable.
But branches do have names, and you almost always make commits onto branches. We shouldn't give up on tracking branches just because tracking forks is hard.
Suppose I have a branch A with three commits, and then I make another branch B on top of that with another few commits. The Git model essentially says that B consists of all commits that are an ancestor of B that aren't the ancestor of any other branch. But now I go and rebase A somewhere else--and as a result, B suddenly grew several extra commits on its branch because those commits are no longer on branch A. If I want to rebase B on the new A, well, those duplicated commits will cause me some amount of pain, pain that would go away if only git could remember that some of those commits are really just the old version of A.
Not really. Git will recognize commits that produce an identical diff and skip them. Your only pain will be that for each skipped commit, you will see a notification line in the output of your `git rebase`:
warning: skipped previously applied commit <hash>And drawbacks, naturally. Advanced branching/merging workflows become extremely painful if not impossible, which makes mercurial unusable as a "true" DVCS (where everyone maintains a fork of the code and people trade PRs/merges).
How does it differ from an extra line on each commit message saying the branch name, and some options to parse it if desired?
I definitely get annoyed sometimes when I have to put in extra effort to figure out which side of the tree is which.
That's really, really not true. First off, I used the word "inherent", which doesn't mean "immutable"; you can retain all the benefits of mutability if you so desire. Of course, Mercurial historically focused a lot heavier on immutable commits than Git did, but hg eventually found a different path that really makes using git feel antediluvian in comparison.
The second thing to note is that there's no requirement that the 'branch' property of a commit correspond to only one head. Actually, I don't think any of the mercurial repositories I've contributed to ever bothered with branches; there's just simply no need in mercurial to create multiple named branches, the way there is in git.
Finally, mercurial solves the workflow problem in another way, by essentially realizing that there is a dichotomy between public, immutable commits and work-in-progress draft commits. The problem with PRs is that you end up in a situation where you have the unenviable choice between making updates with 'address fixes' commits that pollute history or rebases that risk making comments go into the ether (especially on GitHub). You might have extra squashes or rebases that make PRs that depend on other PRs painful. Mercurial instead makes a rebase or other history edit simply mark the old commit as dead and link to the new version, so that any other commits that depend on it can know how to be updated to the new version. And this information is spread to anyone who pulls from your repo, but need not be retained when pushed to anyone who didn't know about the old dead versions!
1. https://git-scm.com/docs/scalar
2. https://github.blog/2022-10-13-the-story-of-scalar/
3. https://devblogs.microsoft.com/devops/introducing-scalar/
4. https://devblogs.microsoft.com/bharry/the-largest-git-repo-o...
There's a bunch of related features they added to Git to achieve scalability without virtualization, including the Scalar daemon which does background monitoring and optimization. Those are all useful and Scalar is a welcome addition. But the need for a virtual filesystem layer for large-scale repositories is still a very real one. There are also some limitations with Git's existing solutions that aren't ideal; for example Git's partial clones are great but IIRC can only be used as a "cone" applied to the original filesystem hierarchy. More generalized designs would allow mapping arbitrary paths in the original repository to any other path in the virtual checkout, and synchronizing between them. Tools like Josh can do this today with existing Git repositories[1].
The Git for Windows that was referenced isn't even that big at 300GB, either. That's well within the realm of single machine stuff. Game studios regularly have repositories that exist at multi-terabyte size, and they have also converged on similar virtualization solutions. For example, Destiny 2 uses a "virtual file synchronization" layer called VirtualSync[2] that reduced the working size of their checkouts by over 98%, multiple terabytes of savings per person. And in a twist of fate, VirtualSync was implemented thanks to a feature called "ProjFS" that Microsoft added to Windows... which was motivated originally by the Git VFS for Windows they abandoned!
[1] https://github.com/josh-project/josh
[2] https://www.gdcvault.com/play/1027699/Virtual-Sync-Terabytes...
But most repositories are not that big so this is hardly an issue for most people. Personally, the system I'm most optimistic about in 2024 is Jujutsu. I've been using it full time with Git repos for several months and it's overall been a delight.
Oh that's quite helpful. I was worried about how lossy the migration would be.
Edit: yes, https://docs.gitlab.com/ee/architecture/blueprints/activity_...
No matter what people might say, I think this stuff matters for contributors and users who might be looking at your project, and git/github is the typical expectation. This is likely the right decision, as they are now ubiquitous.
What I really want is something that will let me use the interface of hg's power tools (revsets, phases, changeset evolution) on an existing git repository.
It isn't a 1-to-1 hg clone, either. But tools like revsets are there, "anonymous branching", log templates, cset evolution is "built in" to the design, etc. There is no concept of phases, we might think about adding that, but there is a concept of immutable commits, so you don't overwrite public ones. The default output is designed to be succinct and beautiful, so it remains relevant on high-traffic repositories with lots of work-in-progress patches, and many developers.
It also has many novel features that make it stand out, like the working-copy-commit. We care a lot about performance and usability; to the extent performance is bad, some of it comes down to piggybacking on Git's data model and existing performance issues. Give it a shot. I think you might be pleasantly surprised.
Disclosure: I am a developer of Jujutsu. I do it in my spare time.
P.S: You might alternatively like Sapling, from Meta. It actually is a fork of Mercurial (you can see it in the UX and features) but is very different now; in particular it also uses the Git data model for the storage layer, so it works with GitHub. It will probably feel more familiar than Jujutsu at first. And it has some absolutely amazing features like `sl web` we can't match yet. https://sapling-scm.com/
Everybody wins.
I don't think PyPy gains anything from this, not even a reduction in the annoying messages that have been psychologically torturing the maintainers. If anything, you're just opening yourself up to more common and frequent low-investment pestering.
SEO : WWW structure :: gravity : orbital mechanics