Anu: A sound, distributed version control systema
anu.dev
anu.dev
Here are some attempts at short descriptions of what I think this aims to achieve:
- pijul 1.0. A stable repo format with performance problems resolved and a good foundation for further work
- darcs except the algorithm is more convincingly correct and merges don’t take exponential time
- A version control system where certain things behave in reasonable ways avoiding potential strange behaviour, eg it doesn’t matter what order you merge things in, you always get the same result.
- A version control system that provides a good user interface to humans, a simple mental model, and asymptotically good performance.
Git isn't perfect, but I've been using version control since Apple Projector (in the late 1980s), and Git has done the best for me. I've been using it for many years.
I don't miss Projector one tiny bit.
VSS (Visual SourceSafe) was a dog. It was direct file-based, and server connections would get very busy. It was the old-fashioned kind, with the need to check out files.
But it had one very cool feature: You could create "aliases" of repo components; essentially creating a virtual repo that pointed into several other repos, taking just a couple of files from each.
I could see how that would be a technical nightmare to implement, but I like it a lot more than "the whole kit & kaboodle" approach that Git takes.
I also used Perforce for many years. It was a robust and dependable system, but had that need to check out files to work on them, and that drove me nuts.
I like Git, because it is "team-friendly," and has a really light touch. It encourages many small checkins, which is how I think I should usually work.
I wish it handled big files better, but that's not really a big deal to me. I think this might be why Perforce is still preferred for game development (their asset libraries get big).
Oh, also Submodules suck like a supermassive, galaxy-core black hole.
Never heard of Apple Projector before. I've always been interested in the history of version control systems, so I would like to learn more about it. But when I search for the term, almost all I find is stuff about using projectors with Macs/iPhones/iPads/etc. Can anyone point to any information sources on it?
UPDATE: It gets a brief mention in the "Legacy" section of the MPW Wikipedia page: https://en.wikipedia.org/wiki/Macintosh_Programmer%27s_Works...
It's obscure for a reason. It was the best back then, but was still a nasty bear.
[0] https://vintageapple.org/macprogramming/pdf/Programmers_Guid...
https://www.mactech.com/1997/12/20/md1-cwprojector-1-0-relea...
In later versions of MPW, Projector was split off as a separate executable, SourceServer. Searching for MPW SourceServer finds a number of hits.
They suck like a chainsaw would sucks for cutting twigs. But they are very effective (together with symlinks) when you need to stitch an amalgamation of multiple repositories (which themselves could be stitched from multiple repositories). That's not to say you don't loose a finger or two now and again.
They are just abysmally badly implemented. They barely work at all and frequently get your repo into nonsensical states for NO REASON other than that the tooling is utterly worthless.
Mercurial has subrepos that offer conceptually the exact same functionality, but is actually implemented in a way that works WITH you rather than against you, and that will not constantly break.
Git submodules are just a completely unfinished feature that is nearly unusable.
It's a long chain of submodules, and making tweaks to the lowest layer means a lot of pulls. I was able to semi-automate it with a couple of batch files.
I didn't use Composer, because I figured that this was a project that would remain fairly static, and submodules are "native" (I am always a bit leery about depending on third-party package managers).
Which has been the case, except that I'm writing an app that uses it as a backend, so I've had to make a few changes lately.
Perforce - I used it too but after their upgrade killed the repo I got rid of it. I understand that it could've been my fault but still ...
After that it was Git
Edit: crates.io says gpl2
But there are scenarios in which this avoids tedious human merges. Consider that I'm applying a series of patches which make changes in a file and later on walk those changes back, and run into a merge conflict there. In Git, I could squash those changes to avoid dealing with the conflict, but then I've lost history. I could apply the changes, skipping the relevant patches, but if those patches still contained useful work elsewhere, then I'd have to go in and resolve those problems manually.
In contrast, this same scenario in Pijul and Anu would just trivially work in a way that didn't produce conflicts. I would apply the sequence of patches, and one patch would produce a conflict… but because they can keep doing work in the presence of conflicts, then they could keep applying subsequent patches and apply the patches which walk back the changes, and in that resolve the conflict automatically, but unlike the Git approach where I flattened the changes first, I would still have the full commit history associated with that sequence.
Now, that doesn't mean that Pijul or Anu will automatically fix all merges. If you have two separate code edits to reconcile, you might still need a human in the loop to reconcile them. But the fact that they can keep making changes in the presence of conflicts allows them to avoid a certain kind of "busywork" that comes with managing git history.
I don't think it is really true that git requires you to fix conflicts before continuing. There are strategies that can let you emulate deferred merging. I rarely let merging hold me back.
If I don't have time to merge into master, just do git push server master:synced/master.
I port/test a project to a new OS. I run into a bunch of issues, like linkers, environments, etc. that are broken that I need to fix. I try and make clean commits tackling one issue at a time, so we get 1 commit for the linker issues, 1 for the docs, 1 for the environment, etc. eventually I have 5 commits of fixes.
These fixes are all orthogonal, so ideally I want to make separate PRs and separate reviews for them. But locally they're all tied together (in chronological order), since I need all of them there to continue development.
In git I can either open 1 big PR including all fixes at once (annoying) or I can make a PR for one commit, wait until it's merged, then PR the next, etc. The only way to get them nicely separate is if I take all those 5 commits, rebase each of them onto master in its own branch and PR those 5 branches. But that ruins my ability to work locally.
The associativity of (non-conflicting) patches in a patch based VCS like Anu/Pijul means that this "ordering" of patches doesn't exist. I have 5 patches you don't, I can PR each of them independently without needing to manipulate history or anything, because fundamentally the patches aren't related and therefore there's no reason for an ordering like the git DAG forces upon you.
Of course this is a fairly niche (but hopefully concrete enough) example of how this enforced ordering of commits can actively harm workflows/collaboration.
Once the PR (possibly with some changes still) is approved and merged, I would rebase my "work" branch on top of the updated master. Since it was an orthogonal commit, you'll typically see that the merged work just disappears from your branch during the rebase. If it really was orthogonal work, you'll not have any conflicts.
This is an approach I've used whenever I need to work on something that depends on other work that was not merged yet. You have a work branch that includes all the work, since you depend on it, but you want to offer smaller pieces as PR to make the review process more efficient.
(He means on his private computer)
That is incorrect and is not a conflict.
> You can think of this as subtracting older from yours and adding the result to mine, or as merging into mine the changes that would turn older into yours.
The point stands: taking history in account is unconvincing, as it assumes intent where none is explicitly given. In the example, if the top branch had a single commit, AB => ABGAB, then there is no way to reliably infer that A 'intended' [+ABG]AB, vs. AB[+GAB]. Even with the given history, maybe the author of the top branch really intended AB => AB[+GAB], but made an error, corrected in the 2nd commit. The problem is fundamentally ambiguous. We can argue which of the diff3 or anu/pijul heuristics are better. In practice I suspect the difference is not that large.
For completeness, here's a guess of what diff3 does, assuming merging top into bottom, following https://blog.jcoglan.com/2017/05/08/merging-with-diff3.
[bottom] => A [+X] B
[top] => AB [+GAB]
B O T
A A A
+X
B B B
+G
+A
+B
Resulting in A[+X]B[+GAB] without conflicts.Then that example you link means that you should stop using 3-way merge / git / svn / mercurial.
Git definitely could be much smart about merge conflicts, but it's a hard research problem so I'm not surprised it isn't.
If I were the Anu people, I would focus on having a seamless compatibility layer that could manage Git <-> Anu repositories (there are undoubtedly many headaches that would occur synchronizing the two different models). This would allow developers to silently interact with ongoing git repos using the "better" tool. Getting wholesale migration to a new platform seems a significant challenge, but allowing developers to slowly build mind share with an improved workflow would be possible.
Disclaimer: I hate git.
That role is filled by Subversion, CVS, VSS etc. with their tragic anti-features.
With the latest kurfuffle at Github, I've started moving to fossil. Having everything, wiki, pull requests, etc. as part of the repo is looking like a good move.
Why let yet another corporation have control over something they should have never been given?
It would be equally as valid to self-host gitea/gogs/sourcehut/gitlab and/or an issue tracker of your choice, which arguably is preferable to adopting a completely different tool over what is a provider issue.
Going against git is an atrocious user interface (if it were good then [1] would be neither funny nor sad). Most people just memorise a few commands and if they stop working they transfer their changes elsewhere, delete the repo, and start again. Sometimes a team will have a “git expert” who has merely memorised a few more commands and is better able to get a repo out of a broken state. Git fails badly at an important for a developer tool: largely getting out of the way.
In fact I often work with git using the tools that come with it: git on the commandline, gitk for the visualization of the history and 'git gui' for committing work. They might look outdated, but they work really well and are really fast.
There are a lot of modern graphical tools to work with git repos, but some of them introduce their own vocabulary in an attempt to make using git 'simpler', but this just ends up making everything more confusing.
My impression is that a lot of people want powerful tools, but do not want to invest the time to learn how to handle them. The Pro Git book is available to read for free online and after reading the first 3 chapters, you should know about the most important things for day-to-day use. Some people would do anything to avoid reading the documentation: they'd rather spend a whole day checking different guis that make things look easy and familiar instead of spending that time reading the documentation.
I've definitely pulled out the BFG here and there to clean up credentials but that's an issue in any VCS.
Maybe I'm biased because I'm "better able to get a repo out of a broken state", but for the record it's definitely not because I've "memorised a few more commands".
I'm by no means a git expert (I've actually just recently learned about bisect for instance), but I have never in my entire career been in a state where I'd just delete the repo and recreate it from scratch.
I've only ever used a handful of commands, the most advanced of which could be probably considered `reflog` when I wanted to revert some changes; or `rebase` (because strictly speaking, it is more complex than merge I guess), but I never ran a command I did not understand or had to memorize.
I actually do share the sentiment about the tool getting out of your way, and my knee-jerk reaction to learning about git internals is just repulsion, because you're right! I'm not there to tinker around with version control, I'm there to solve problems. That said, I've never felt like Git got in my way.
The one and only time I messed up a repo beyond repair was when I deleted some git pack files while trying to delete some binary files from the git history. This is known as user error.
In my day to day use I find that I rarely have to venture beyond rebase, bisect, reflog, cherry-pick, and the standard commands.
Whether self-hosted git or hosting on Github, your issue trackers and such are typically separate from your main repository. Most platforms offer wikis as a side-by-side repository so that should be easy to move, but the rest is at the whims of the platform.
The GP is claiming they moved to fossil because the one repository contains all of this data.
I haven't followed Fossil, so hearing that it includes things like a wiki is news to me.
As a nice bonus Fossil is a single executable/binary file you can drop anywhere and can act as both the CLI for working with the repository and as the web backend with a bunch of ways to access it including CGI, it's own web server or even as a fake script parser (you can upload the linux binary to any shared host that supports custom script parsers -many do- and use a "script" with a shebang that calls the binary with the path to the repository file, thus allowing you to use Fossil with shared hosting services that do not even know about it).
I'm pretty sure github doesn't control git.
https://discourse.pijul.org/t/is-this-project-still-active-y...
Pijul was always advertised as experimental, and Anu is the result of that experiment.
So it should actually be "safe" to buy in to now? If I put my project into Anu, I shouldn't get stranded in 5 years?
The biggest contribution I got was from someone who rediscovered it independently after asking me about it, and didn't even care to respect the license.
Moreover, the formats of an experimental tool, especially when it is based on new math, need to change frequently. Every single time we've done it in the past, we heard weird comments here, on Reddit and Twitter that our theory would never work because the implementation was not there yet.
There were also unfortunate professional choices I don't want to comment on, which forbade me to work on Pijul other than sometimes on the weekends. I ended up resigning in July 2020, and have worked on the new Pijul more or less 100% since then.
We all know why
"It is based on changes rather than snapshots"
Well, every VCS I'm familiar with is based on changes/deltas. I assume that these terms have specific meanings here that I'm not familiar with, but it manages to sound like the author has never heard of git or Mercurial.
[0] https://git-scm.com/book/en/v2/Git-Internals-Git-Objects#_gi... [1] https://git-scm.com/book/en/v2/Git-Internals-Git-Objects#_tr...
I don't know how general they could be, but darcs can have different patch types. However, the only extra one implemented, as far as I know, is token replacement.
Is that where the innovation is here, Darcs style patch sets?
It's surprising to hear that svn tried to be clever with merges, because its merge support was no better than CVS until svn 1.5, which was released some years after git. (This is one of the reasons I went straight from CVS to git.)
Combined with the fact that merge and commit were two separate steps in the workflow with manual conflict resolution in-between, this lead to sometimes severe usability issues. The merge would adjust the mergeinfo properties along with the files. Users would sometimes go and do svn revert on some of those files as part of conflict resolution. This would also silently reset the mergeinfo properties that were just updated. The current merge would still work out OK. But future merges involving the current branch or its descendants would get horribly mangled because SVN ends up applying sets of changes that are out of sync with the actual state of the files involved. That's part of what gave SVN its bad reputation.
Maybe we're talking about different versions, though. I used to use SVN 1.6+ IIRC.
Unless a merge is fairly straight forward, I often use --no-commit and only select the files I want to merge, and defer a complete merge for later. It is so easy to make intermediate branches, or just pick what I want from one branch to another.
I haven't used Darcs or Pijul, but I feel like trying to "solve" merging and source code isn't fully possible.
Not if the functionality you want a patch-based system, which is just so much easier to use. After 30 years or so experience of using revision control systems with distributed projects, I can't see the appeal of git.
This is returning "Not found" for me.
There’s a frightening amount of stuff to write, and it will take me a while to explain everything.
My current plan is, I’m actually releasing it as I’m writing this answer. There is almost no documentation, but I’ll write it one page at a time in the next few days. I’ll also write a blog post tomorrow to explain what I’ve been doing.
Oh, and after talking to Florent, we have finally decided to change the name. More on that in my blog post tomorrow.
https://discourse.pijul.org/t/is-this-project-still-active-y...Can somebody give me a real world situation in which Anu would work better than git?
The underlying model sounds like a big improvement, but I still can't map it to the benefit that I as a user would have.
The rebranding seems to be for several reasons including:
- Breaking changes from pijul releases
- "Quicker" writing of the command "anu" vs "pijul" (they also mentioned on Birdsite that anu is easier and faster to type on Dvorak)
- Easier to spell out and pronounce vs pijul for non-latin-language-speaking users.
Most real complaints about git are around scalability of giant monorepos, and a lot of work has gone into various solutions.
The secondary complaints about usability seem to be papered over by popularity, and of course the relevant xkcd: https://xkcd.com/1597/
For me, I'm excited about better, more rigorous merging and being able to cherry-pick & rollback changes without causing conflicts later on. (Cherry-pick in Git makes a new, unrelated commit, so merging with the branch you cherry-picked from can often cause merge conflicts, etc)
In general, tracking and working with actual dependence between patches seems to open up more workflows, and less hacky ones.
e.g. giant monorepos weren't a problem even for ancient centralized VCs because these were file based, and if you wanted to work on a part of the tree it didn't matter - just push and pull the part you care about. But git has a data structure that forces everything in a single tree, so you have to use hacks (submodules etc.).
Same thing for the much of the UX and the other complaints (changeset/patch model). When you get down to it, the data structure is behind 90% of difficulties with git.
This is due to the lack of a solid theory to match the intuition. The "git way" of trying to match each use case is to add a new command for each new use case.
I don't think you can fix "just the UX" without fixing the underlying algorithms.
To me, git's problem is the exact opposite - the commands are based on how the underlying technology works, not on the use case for the user.
The best example is 'reset'. I understand git's internals reasonably well so I know the reason why, but it's not obvious that you need the same command to "un-add" a file, or to wind history back two commits.
A talk introducing the issues with git and how gitless improves upon them https://www.youtube.com/watch?v=31XZYMjg93o
The patch model is simpler IMO
And Anu is apparently even sounder and faster?
https://pijul.org/manual/why_pijul.html#pijul-for-darcs-user...
git was built as a tool for completely distributed source versioning, but most of us are using it in a centralized way. It's nice to be able to work offline, but when we need to synhcronize there's always a huge dance of fetching first, see if it has moved, merge/rebase, etc... git is good at storing what we did, but it doesn't help at all at saving what we are _doing_: all changes to the working directory are ephemeral, like files stored in ramfs. When working on public repositories, you can't push a branch prefixed with your name; you have to fork the whole project _and_ push a branch before you can start interacting. Instead of having one server and a client, you now have 1 central server, 1 other server that only _you_ can access and will in practice contain 1 branch, and will be abandoned as soon as you're tired of it, and a client. Rights can't be managed at the branch level, so I'm just going to copy-paste the whole thing from the beginning of history and give it to you.
What I would like to see in a VCS:
- There is one central place where people coordinate - There is exactly one commit associated to a branch, and that association is the same on all machines at the same time (I don't want to git fetch) - If you want to do changes to a branch, you do a sub-branch - That sub-branch, along with your local changes in or out of the staging area, is synchronized to the server. If authorized, other clients can have a view of those as well
It seems it already exists with fossil (https://fossil-scm.org/home/doc/trunk/www/concepts.wiki#work...) and with older SCMs, although older SCMs are plagued with the locking problem.
In a way the work that is done to handle giant monorepos is helping git move in this direction: all branches are automatically synchronized, and the vision with this kind of repo is that it's ok to commit often, even in small batches. But it's not quite there yet. I've read an account of how things are done in Google (https://cacm.acm.org/magazines/2016/7/204032-why-google-stor...) and it's closer to my dream system.
I believe this is more of a issue with GitHub then with git itself.
Still, it's not enough in my taste because I want all content to be synchronized everywhere
I believe that you need to understand the semantics of the code to truly do what you are trying to do well, and for all other cases the snapshot model is more than good enough and given how we structure and modify code, it works out really well in practice. Code dealing with a single aspect should and almost always is co-located, so to get a conflict of intention in a merge is very rare. There are other human aspects like code ownership and collaborating teams which makes the issue even less of a problem.
To do something similar with more structured files, one must find the corresponding idea to “a list of lines”, and this must work in a good way (e.g. changes like x -> (x); [a; b] -> [a] foo [b]; [[p, q], [r, s]] -> [p, q, r, s] must in some sense be natural operations in your structure (and diffs need to be reasonably easy to compute)). And of course it still needs to work in a sane way for unstructured data in big comments. Therefore I don’t agree that Anu would be easily generalised to this.
I think this is basically impossible to do for situations where you want to capture all the structure (such that a patch to rename something merges well with other patches). I think it’s likely extremely hard for a part way solution.
Finally I’m not convinced that the change would be that useful. Much of the structure of computer programs is implicit in the scoping rules in such a way that the “move blocks around” changes that line-based VCSes often struggle with will still be invalid with structural diffs.
Google doesn’t need a different representation where all push outs exist because they rely on a centralised server, low latency, and arbitrarily choosing how to resolve conflicts. In a DVCS, you can rely on none of these.
I still claim that the reason OT works well with google docs is that it can rely on a centralised server, low latency and tie breaking.
Tie breaking means one doesn’t need to worry about representations of conflicts (and allowing changes to merge in sound ways) which is in some sense the main thing pijul does.
Low latency means that users are able to cope with the tie breaking rules doing the wrong thing
A centralised server means that there is less need for the merges to work in the sound way that pijul aims to make them work.
Therefore I put it to you that google docs is neither an example of the same theory that pijul is based on not evidence that OT would work for some kind of well-behaved structure-aware DVCS.
> Anu is written in Rust, and can be installed by first installing Rust, and then...
Yeah, I'm not installing an entire language just to use your tool. I don't need to install a C or C++ compiler to run Photoshop or Microsoft Word. Why do I need to install a compiler, libraries, etc. just to try out your tool? No thanks.
Doing it from source also has the advantage that the package maintainers can customize the build process so that it works better with their system. Photoshoto and MS Word are closed source and proprietary, which creates issues if you want to package them.