What comes after Git?
twitter.com
twitter.com
What's stopping us from just using those words to describe operations with Git?
I mean, what Git is, is a data structure with a set of operations on it, and the design of that is pretty fundamental at this point. We'd need to see something more novel[1].
(Sidenote: Git almost seemed an obvious trajectory for version control given how problematic file-based approaches have been. I've used far too many version control systems, including "ancient" awful ones like Harvest[2], and it just felt inevitable that we'd one day rely on an "always-branching" model.)
[1]: I did think of something more novel: semantic version control (https://news.ycombinator.com/item?id=25537455)
[2]: https://en.wikipedia.org/wiki/CA_Harvest_Software_Change_Man...
(Unless I guess we're talking forks of MySQL.)
I don't know if maybe you have a good example...?
The protocol has become a bit smarter, by allowing there to be additional options passed for different communications options, and there is a slow but steady progress of moving from sha1 to sha256 for the hash function, but that isn’t rolled out yet completely.
Different tools have been added over time but the internals are pretty much the same as they’ve always been.
Of that was the case then how come Git reigns supreme in spite of the lack of a polished UI/UX?
The huge existing body of documentation, tutorials, StackOverflow questions, etc. that use the confusing standard Git terminology.
I feel like making collaboration less painful has been the greatest innovation driver in the last few decades, and has paid off absolutely enormously. Imagine where a git/Github for other disciplines could take us. Given software developers build such tools, and given software developers overwhelmingly know how neat git really is, it's actually kinda strange that there's still such a dearth of good solutions in this space.
I would love to collaborate but when I start looking into git the amount of stuff I don't fully understand keeps growing. I should want something in my work flow that I don't understand? If it stops working just like that, the way everything seems to, I'm suppose to fix something that I don't understand?
When trying to figure out how it worked I read discussions from truly experienced users who decades into the process apparently still had to learn now to unstuck themselves.
I'm not saying its bad at what it does. Its much worse, I have no idea what the point is. It seems tailored to manage top of the developer pyramid issues that I imagine to be hard enough to justify its existence.
In applications I've seen copies of data get modified in different places then merged back together again. Its a horror movie!
As a lone developer, it lets you easily snapshot the state of your code at appropriate points and then compare those to the current state to find out what you broke. Of course, you can also do this by making copies manually, but that gets complicated if the program consists of more than file.
The real benefit is when you have more than one developer working on the program. If you make some changes on your machine, and I make some changes on my machine, and when we're done we want a single unified version of the program (with both your changes and mine) in order to deploy it, what should that look like? If all we can see is that line 123 in some file is different in your version and my version, how do we know which version of that line we want?
Git keeps track of (in the form of a graph structure) the fact than your version 1 is a descendant of our original version, your version 2 descends from your version 1, my version 1 descends from the original version, and so on ... and when we merge our versions, we get a final version that descends both from your latest version and mine.
You could do all this bookkeeping manually and figure out what to diff against what, but git gives you tools to work with.
...and manual copies don't offer you the great automations that git provides you for free, like being able to tell when exactly each line has been changed last time (so you can check the whole context of how and why that change was made) or perform a binary search of a revision that introduced a bug. Even when working alone, these tools are invaluable.
We don't expect the average person to understand the tools and workflow of a mechanical engineer working with SolidWorks PDM - we shouldn't expect the average person to understand git (which is basically the software equivalent). I've never disagreed with the arguments that git is confusing or difficult to learn for non-software-developets, but I've never really understood why that's a problem.
Everyday folks aren't expected to know how their car works, much less tinker with it. Not the enthusiasts, who will find a way to grok the full specs of cars and tinker even if you tell them not to.
Normal everyday folks should be content leaving the gory details to their mechanics. I'd wager it to be the same with Git, as an everyday tool of software engineers, rather than that of layperson. This is not gatekeeping; it's a fact that Git is not your usual point-and-click stuff and UIs that attempt to do so can only simplify it so much. Real enthusiasts can always learn how to use Git if they want to, there's no need to make it simpler or easier for the 'casuals'.
The fact that people even have trouble with self-descriptive software like Microsoft Word - which is probably, in terms of UI, one of the easiest, most accessible UIs ever with years of research and iteration - speaks to the fact that not everyone will 'get' Git, nor sho uld there be a need to please or try to reach to everyone with a software that has as many complicated features as it does for the multitude of engineering problems it tries to solve.
If that were actually the case, then the dominant programming language today would be COBOL, rather than anything we have now.
The difference between my example vs. what you defined, is that hype and/or popularity determine what people use for _new_ projects. But that doesn't really assess how dominant a language has been in the past (or even in production currently).
Java is another dominant language that you might not use for a new project. Same with C.
That definition biases towards verbose languages with lots of headers and boilerplate code (e.g. C++, Java)
I would very much consider using Java for a new project, same with C(1) if the project is in "real time" (I think that C++ hides too much things under the rug which can create latency issue).
1: Zig isn't ready yet to be a replacement
Those are the usecases that Git already handles. Each and every single one of them. Somehow you conspicuously left branching out of it, which is the main value added by any version control system along with merging. Thus I really don't see the point of arguing that Git will be replaced by a tool that does exactly what Git has been doing by design since it was released.
You really want to record, mapped X to Y. Not 500 files changed in 4218 places.
I imagine this could be achieved by using more intuitive nouns for subcommands (reflog, rebase) and giving informative warning messages on the command line when you are about to screw something up irrevocably.
Eventually, one would hope that some of the more effective features of such a wrapper can find their way back to the git project itself.
personally i think putting out sites like that is more helpfull than dumbing down git
git config --global pull.rebase true
You should also enable `rerere`[0]: git config --global rerere.enabled true
[0] https://git-scm.com/docs/git-rerereThere is no reason fetching and merging have to be connected, so perhaps the most prominent command (as a counterpart to "push") shouldn't be this combined operation.
1. Work on my branch. 2. Checkout to master (with `git stash` if needed) 3. `git pull` 4. Do whatever I wanted to do with the updated branch (switch back and rebase, or test something, or start a new branch etc.)
In a code-reviewed, merge request-oriented workflow there's little point in splitting `git pull` into separate commands.
It's up there with stashes (which actually are just temporary commits ahead of the head anyway) and staging as things which shouldn't be there by default.
Git really likes making things complicated for newcomers.
Darcs failed to reach mass adoption, and lost against Git back in the mid-2000s, which for me, at the time, felt like a big step backwards (and in other ways a step forward, since Darcs had some icky bugs and performance issues). Some 15 years later it's amazing that we are still stuck with Git.
What makes Darcs different is that it does not need to organize history sequentially. It knows what commits depend on other commits, and a branch is simply a "sea of patches" whose order is inferred. This means that you can share commits across branches without conflicts (unless the branches actually conflict, of course). Darcs' equivalent to "git cherry-pick" simply pulls along with it any commits that your cherry-pick depend on.
Darcs had flaws that prevented it from succeeding, and the fact that it was written in Haskell didn't help attracting developers. Darcs' philosophy lives on in Pijul, which also looks promising, but unfortunately seems doomed to the same fate of remaining a niche tool. Git ate the world thanks to GitHub, and the next VCS needs a similar kind of killer app that justifies the switch.
I think the next VCS needs to offer a value proposition other than just "better than Git", though. The easiest way to accomplish this would be to piggyback on something popular.
For example, if Microsoft started offering a new VCS system built into Visual Studio Code that were enabled for free by default and had awesome built-in integration (VSCode's current VCS integration is nearly useless, IMHO), then I think it could reasonably compete with Git, though it might require competing with Github. (Of course, Microsoft could do this with Github.)
Even with this, I don't think a better Darcs would suffice. For example, one thing that Git handles poorly is large files, and it doesn't do well with large repos, either. I think we'd need something like Darcs paired with awesome scalability (Perforce-style central server with partial client-side checkouts) combined with support for arbitrarily large files.
To be fair though, it's a source tracker which are usually small-ish text files. It's not really meant to be a Digital Asset Manager, say.
Git can do a shallow clone as well for very large repos. Much of that deep information becomes probably useless over time, unless you're doing code archaeology.
Also, there's no particular reason Git couldn't be modified to work well with large files.
Apologies for the late response. Yes. I absolutely agree with that. The trick though is that when changes happen to a large file, what's the best way to do a diff? What if it's bzip2 or encrypted?
The other issue is that stagnant large files in your history need to be pruned. I'm not sure how to do that in git right now other than rewrite history.
Is that property relevant, though? I mean, although it does not need commits to be sequential, well... They already are, and there is no way around that. And sharing commits across branches, whether it's in the form of cherry-picked commits or pure old diff/patch workflows, is something that Git already handles.
Additionally, is there any workflow that does require any of those properties/features?
Meanwhile, how about Darcs' computational cost?
With Darcs, all merging can be thought of as rebasing, except histories remain compatible and mergeable. With Git, on the other hand, rebasing the same commits multiple times result in the same conflicts that need to be fixed every time, which is why "rerere" exists. Darcs does not need "rebase" or "rerere" or any of the other tools for complicated history surgery.
As for workflows that "require any of those properties": There's wide agreement that Git is overly complex (exhibit A: hundreds of HN posts about learning to use Git) and is especially daunting for beginners, compared to easier tools like Subversion and Mercurial that provide nicer UIs. Darcs brings modern version control to users without incurring the complexities of Git.
It's not perfect — it had/has corner cases where merging became computationally expensive, though I hear there are new developments to fix this — which is why I said the future isn't Darcs.
That's the part I don't get from your argument. Git has rebase. It works. So where exactly do you see the value added by Darcs? I mean, it's not as if users can only merge or rebase if they use Darcs, right?
> As for workflows that "require any of those properties": There's wide agreement that Git is overly complex (...)
No, there isn't. In fact, that statement is patently false. git is the simplest and most straight-forward version control system I ever used. Things just work out of the box, they work well, and they work exceptionally fast. Branching is free, merging is free, tagging is free, it's trivial to checkout, it's trivial to setup a shared remote repository, and it's even trivial to setup a remote repo on a USB pen. You don't need to install servers to get it to work locally or remotely, you don't need to bother with extensions to get basic functionality working, and you don't need to use absurd ad-hoc patterns to get basic features such as branching.
So is there actually any objective point that can be made in favour of Darcs? I mean, being fast and usable surely is not it. So, what real-world value does Darcs offer?
With Darcs, you can. You can merge/cherry-pick any history without incurring other conflicts than "real" conflicts where people have worked on the same lines of text. This eliminates a large group of commands Git has to provide ("rebase", "cherry-pick", "rerere", "reset", etc.) because the entire problem space doesn't even exist. Darcs is utterly magical that way.
> No, there isn't.
I have to disagree there. There's a huge corpus of material showing that people struggle with Git.
For example, a cursory search of HN shows many, many threads that devolve into discussions about how unnecessarily complicated and complex Git is. You can also go on Stack Overflow or other forums and look at all the users whose repositories are stuck in a state they don't understand. I suspect no other VCS has forced so many users to give up and wipe their local repo to start over again. Git's UI has gotten better over the years since I started using it around 2007, but it's not great at explaining repository state.
Many less-than-technical people — such as designers — have been forced into using Git as part of their work. Subversion was extremely popular in this group, and the transition to Git was slow precisely because Subversion offered simplicity where Git didn't.
Anecdotally, I work with a bunch of highly competent, highly technical colleagues — some of whom I consider to be better developers than myself — who still sometimes struggle working with Git. Most developers don't want, I think, to be "Git experts"; they just want to get stuff done. Most users, myself included, arguably use just a tiny subset of Git commands that deal with the daily pull/push/commit workflow. When I joined my current company, people weren't even practicing clean histories, but using "git merge". But once you get into rebasing (which many people do not do) and diverging histories across multiple levels of branches, you open a whole can of worms that Git is not good at dealing with. Rerere, while a huge time-saver, doesn't solve everything in such complicated scenarios.
Perhaps you're perfectly fine with all of this, but in that case you are not representative of the average user, in my opinion.
> git is the simplest and most straight-forward version control system I ever used
Since I started out, I've used RCS, Visual SourceSafe, CVS, Subversion, Perforce, GNU Arch, Darcs, and Mercurial. The only tool more complex than Git was GNU Arch (though, admittedly, required jumping through some ridiculous hoops to work around its lack of native branching support), but Git is without question the most complicated and complex that has been in mainstream use.
So I've used mercurial and git in production. The one thing that turned me off on mercurial was that branching information was stored as part of the commit. This led sometimes to weird conflicts where the branch list needed to be merged. Granted that's been several years ago, so it might be different now, I don't know, and with git I have no need to.
I definitely appreciate branching data being meta data outside of the branch.
Also Darcs had severe issues maybe a decade ago where merging would appear to cause infinite locking. It wasn't a coding bug, it was due to their originally flawed theory of patches. So while it may be awesome, the world has definitely moved on.
It's worth reading this from pijul:
> Did you solve the “exponential merge problem” darcs has?
> Yes, we solved the exponential merge problem. The only caveat is that Pijul does not (yet) have an equivalent of darcs replace. In other words, Pijul works in polynomial time for all patches that systems other than darcs know of. We’ve not yet thought all the theory of this through, but it might be added in the future.
So good, but polynomial time can still be large beasts computationally.
Other than that, I like not having to deal with the additional complication of Git's staging area in Mercurial (anything I personally did with the staging area I can replicate much better with Mercurial interactive commit feature), plus Mercurial's evolve extension, which turns commit amending, rebasing and all other kinds of history editing into non-destructive operations.
I would also like to have the following workflow (in functional language like Haskell, in more imperative language it would be trickier):
Before I start working on a code change, I mark the functions and types that are to be modified. A tool then calculates a "change boundary", a set of function calls that potentially have a different semantics in the new code. After I am happy with the boundary, I will start working on a replacement code inside the boundary.
The old code will still be available alongside to run and inspect during the whole development process. An automatic test suite generator will run the tests on the old code and by observing the boundary, it will automatically create a regression test suite for the replacement code.
Once I am done with the new code, and it is tested, I will let the tool replace it in the defined boundary as a new change.
So I will have a guarantee (through types) that I am only changing things that have to be changed, nothing else.
It is a pile of garbage, but it's better than nothing.
Also, deep learning training data often consists of large image files, and can also be considered "source code", and in any case it can be very useful to put these under version control.
And finally it can be useful to put external dependencies as tar-files into your source tree.
For writing tests in a deep learning code base, rather than simply including a native data file (image, CSV, whatever), I've taken to writing a fake data creator class. It always feels like overkill when an alternative solution is including a native data file or two that already exists.
LFS uses some sort of internal filtering and tracking to determine which binary files might have changed. It seems to have trouble deciding if there are actually dirty files that need changed. So you can't just say, "Okay, go find all the binary files that didn't actually get moved to LFS and correct them"
Instead you end up with random moments where you want to commit a single file and git instead detects 1000 png files that it absolutely could not go on without doing something about.
But then the diff is a disaster so good luck understanding that what it is actually mad about is that it wants to move the files into LFS. The only way I finally figured it out was to manually load the object blob and notice one of them was an LFS pointer file.
I personally think git annex handles things more elegantly, but lfs won that battle.
could you elaborate?
* Changing from a subdirectory to a submodule breaks lots of things like git reset and git bisect.
* Having to remember to git module init, update etc. I always have to look up the commands and never remember what the difference is.
* I don't care that there are unlisted files in a submodule, either don't bug me about this in status, or integrate commands in such a way that they work transparently across the main module and submodule.
* Related to the previous: Coordinating a single logical change across submodules involves several manual steps and has plenty of scope to go wrong.
- An edit in .gitmodules
- An edit in .git/config
- A removal in .git/modules (seriously?)
https://stackoverflow.com/questions/1260748/how-do-i-remove-...
And that's the only issue you had with them?
I'd rather believe that submodules are fixed at some point or an alternative solution appears that works much better. Git subtree is around for a while and there's also Git X-Modules https://gitmodules.com which is modules on the Git server.
"Unix philosophy" has always been in the eye of the beholder. I think Perl showed the 1980s Unix philosophy of "small programs, working together", was often overly cumbersome.
Why would it need to be replaced, just because it's 15 years old? I can think of a zillion things that are over 15 years and still very much working.
You don't need to learn emacs beyond C-X 0 and C-X C-B really.
Productivity is higher than with manual git commands.
Now the Jetbrains IDEs have a absolutely terrific Git integration, but if they didn't I could still open Emacs and use Magit.
Many have taken a crack at it over the years, from very early on.
Issue's no porcelain will even replace the built-in one (and what improvements are added to the standard are generally flawed to hell, just look at how messy the brand new `git switch` already is, the git core team simply has no taste when it comes to CLI), and as the creator of a new high-level CLI by the time you'll have good enough feature coverage for it to be useful you'll know the plumbing so well the porcelain will make perfect sense. So you'll abandon the project as offering little value for the maintenance cost. And a few years later somebody else will come in and create their own, and repeat the cycle.
Besides, auxiliary tools aside (I use git shell and a bunch of aliases), the git executable too is still improving. It will be a long time before anything that might replace it reaches both the maturity and the tipping point of having enough extra features to make it worthwhile to consider switching.
It’s also still actively developed so it’s not like some “stale” software we use because of inertia.
Of course, for people that know how to use it. I'd argue that the heavy use of Github combined with the adoption of advanced Git workflows is currently the biggest threat to open source. There's no such thing as "submitting a patch" these days. You're expected to be a master of a project's Github workflow to contribute. Most of us (myself included, with a few exceptions) just don't bother.
With Github, one has to fork the project, run some commands to clone the repo from the fork, make the changes, push them, and then go back to the web interface to create the pull request.
I was fine with that as a maintainer too, on three very active major projects at one time (using a 48Mb Pentium 1). A maintainer's job is partly to make it easy for contributors, not for the maintainer, but tending the contributors will help in the end.
That entirely depends on whether the necessary libraries and headers are installed on the system you're using.
> even something that didn't need a rebuild (config, doc, interpreted source), and sent the change, perhaps using M-x diff-backup
Unless someone uses emacs and gnus or something similar, email clients could end up mangling the diff unless it was sent as an attachment (which would make reviewing the code and posting an inline response more difficult).
Having to set up one's git config or having to fork a repo seem to be accidental complexity is both cases.
The next time they open a PR using that branch, it will have commits that will clutter the branch.
Moreover, I suggest you fork the project and apply your own modifications instead of worrying about a GitHub project administrator to accept your changes
I think it's odd that so many people use a DVCS as part of a development system so dependent on a centralized server for everything else.
Note I love git, but I don't know any big game studios using it. Some indies get by because their games are smaller but even then it can be a pain.
The general point is: Git mostly works fine for what it is, but version control can 1) be easier to use, and 2) serve more purposes than Git currently does.
Imagine a perfect, seamless version control system that does everything you want with minimal effort. Git isn't that, not even if you've put in the time necessary to become an expert.
Often this is true, but fetishing change for its own sake becomes inefficient.
(Over the summer, I made a command-line [https://gitlab.com/jcfields/versions] and a GUI program [https://gitlab.com/jcfields/restore] for accessing this system outside of the standard UI, if anyone's interested, though the binaries are not notarized so the OS will show a scary warning the first time you use them.)
Granted, it's not exactly the same thing since you're not making commits at discrete and meaningful points like in a software version control system and since it works on an individual file level, which might not be ideal for every workflow, but it's still really useful when you realize you need to roll something back to an earlier state.
I think modern versions of Windows use the Volume Shadow Copy service to store backups of files at given snapshot points (such as when your computer runs a scheduled backup or creates a System Restore point), which can be pulled up in the file properties. I use to use NTBackup to do basically the same thing in a more crude way back in the Windows XP days. This isn't as nice as the Mac Versions feature since it requires setting up periodic backups, but it's something.
I'd be curious if any of the free desktops have come up with a simple and user-friendly solution to this. I feel like this is an area where there's a lot of room for improvement, since the majority of users would probably benefit from these systems but most aren't even aware they exist.
It's always been very curious to me that, for so many parts of dev, we re-invent the wheel a thousand times to try to find the "right" configuration, which ultimately leaves you with a very wide array of options for every one job, but for source control we seem to have just kind of shrugged and said "Git's fine."
Version control that is aware of changes not at the file and line level, but at the function level, the module level, the library level, the language-specific construct level.
(Rich Hickey, creator of Clojure, played around with this idea, but I'm not sure where it went.)
Then we’ll find out what its real pain points are, and start working on something new to solve those.
I would guess some kind of answer involving complex branching schemes and/or "octopus merges", but I consider myself a fairly advanced git user and find those too hard to understand. I always stick to simple linear histories with occasional stable branches.
The commit messages could be the ultimate historical record of all the decisions and the "why"s.
Yet, on most projects, saying "I went to read the git log and found out the reason of x" instantly elevates you to a superhero status.
git log -p --follow <file>I've been playing around with a proof-of-concept in Go based on the Restic chunker library: https://github.com/akbarnes/dupver
If you're using git-shell or git-daemon to host your git repositories on a server, then you can't use that alone to also manage transfer/access of stuff that is managed through git lfs.
But beyond that, what’s keeping the server from storing these deltas instead of full versions? Only argument I can think of is that server side corruption of a single file would break the chain of deltas and you’d lose every child.
But surely that can be solved using some sort of redundancy.
I'm making a simple .zip over HTTP solution for my MMO now.
It just has a naming scheme with dates and you send your last .zip date to the server and the server sends a list for you to download.
The good part with dates is that you can consolidate the data retroactively without breaking the system!
A decade ago, something like that, a project I was on switched from SVN to Git. I don't know why, I don't know what problem was being solved.
"Git isn't that hard once you understand the internal model" one of the devs said.
I didn't need to understand the internal model of Subversion, we just checked stuff out and checked it back in. Within a month of adopting Git they had managed to lose a week's work. I guess they didn't understand the internal model. I kept my mouth shut.
For most corporate development teams, there is no problem that Git solves. Corporate development teams don't do distributed development. In a decade of using Git, I have never told another developer to connect to my local Git repo to get a piece of code. I would create a feature branch on the very non-distributed Github or GitLab server and they would go the single source of truth to get that piece of code. Like you would have done with Subversion. Only without the complexity.
In the Subversion days, non-technical people would use the same repo as technical people. Like tech pubs. With Git, they can't.
As I said, I have no idea what problem Git solves for the typical corporate software development organization. But I admit, it is complicated and hard to use, so in the end, as a developer I make more money. The more complicated it is to make software, the more money I make, and that's what's important.
Rock on, Git.
Personally, I think the next big VCS will have the ability to handle all files, including binary ones. (See https://gavinhoward.com/2020/07/decentralizing-the-internet-... .) As a bonus side effect, this would probably give that VCS the ability to handle source code by semantics.
Git is still good for tracking changes in plain text files.
But Git / BTRFS are Linux, Where Hg / ZFS are BSD. Something better could always come along, whether it gets significant adoption has very little to do with its technical capability.
In what way?
They were there at the same time (or before) git, and they don't provide significantly more value than they did then.
While I do think they provide value beyond what git offers (my kingdom for revsets instead of the hell that is gitrevisions(7)!) the "market" rejected them then, it has little reason to switch now.
- Built in better binary/large file support
- Need to be able to checkout subtrees and host monolithic repos (submodules are a pain)
- Better UI
IMO the utter disaster that was gnome3/gtk3 transition is what killed what momentum there was for a mainstream Linux desktop.
i think as long as git keeps evolving and there are no major pitfalls for any programming language: why should it be replaced?