Why SQLite does not use Git (2018)
sqlite.org
sqlite.org
GitHub and GitLab offer nothing comparable. The closest I have found is the network, which is slow to render (unless it is already cached), does not offer nearly as much details, and scarcely works at all on mobile."
This is a valid point. I've always wondered why it seems Github goes out of its way to hide the typical commit tree navigation that so many git clients support (Sourcetree, gitk, tortoisegit, etc).
Eg https://github.com/sindresorhus/delay/network tells me that this work is being maintained actively, but most forks are not merging back. Another one might tell me work on the main fork is stalled, and many users are now doing PRs against a fork of the original.
... across all forks, which makes the UI much more cluttered than it needs to be. What if you are only interested in commits in your own repo?
Github really does seem to lack a clean equivalent of `git log --graph`.
that would be worthless for most people, and for the sqlite authors, who'd surely have a million people pressing the 'fork' button.
it also isn't what anyone in the thread, or post, were asking for.
perhaps you'd be so kind as to provide a link to a page on github which demonstrates the same graph structure?
> why it seems Github goes out of its way to hide the typical commit tree navigation
Wild guess: because those kinds of features requires the back end to execute Git implementation code which can invoke some pretty expensive cases. Imagine that across thousands and thousands of repositories being accessed at once.
They want to steer the user toward doing a clone and doing their own in-depth drilling on the copy.
CGIT caches for better performance; but caches require extra storage and some sure-fire expiry mechanism so it doesn't grow and grow. Caching things for internet rando visitors doesn't seem like a good idea.
Github has the blame view of files, which is much more expensive.
Not the idea of a decentralised system, is it?
Had to give up on it because using two different revision control systems became inconvenient after a while. But I regret the loss of Mercurial, after having used it for a decade or so, more than that of Fossil.
It always bothered me how a user-hostile piece of software like git rapidly took over the world. But then I remembered sh and bash and realized that someone thought that these softwares should be designed the way they were decades after Pascal and Lisp and hundreds of other sensible languages came into being, and they took over the world too.
After almost 7 years on git I have given up on getting git and I stick to what that XKCD comic said - remember some commands and just keep syncing. It felt one has to make it into a masochistic exercise to understand it.
I feel similar about Markdown (though I won't call it user hostile). Because there were and there are better and accessible/uncomplicated alternatives.
I don't know; all of the lightweight markup formats are a mess. AsciiDoc/AsciiDoctor is almost great but they seem to have completely borked escaping of inline punctutation[1]
Restructured Text is very "meh" but is better than markdown because directives at least make it more sanely extensible. The downside is it's missing lots of features in its core becase "you can just use a directive"
Markdown was actually pretty good for its original purpose (write html with fewer html tags), but completely falls down for its current common use case (untrusted user input, where you won't allow the user to fallback on html).
[1] 1971 (https://en.wikipedia.org/wiki/Thompson_shell)
[2] 1979 (https://en.wikipedia.org/wiki/Bourne_shell)
[3] 1978 (https://en.wikipedia.org/wiki/C_shell)
At a conceptual level, so much of the more esoteric parts of git makes little to no sense. E.g., rebasing. I don’t completely understand what it is supposed to achieve.
Even vi, with its own idiosyncrasies, is eventually grokkable once you understand that vi key sequences are an expression of an action you want to take.
Git? Nothing makes sense. There is no connection (that I can nail) between its many features and how we think about and articulate what we want to do with a repo.
It’s success is so strange to me.
Gotta be honest, rebasing is not an "esoteric" part of git.
You would be doing yourself a service to read up on how git and its internals work. If you're a software developer, chances are high that your workplace, or a library your workplace uses will be managed via git and you'll have to use it. Might as well be ready for that, because if you put in just a little bit of effort to understand it, it becomes an easy tool to use.
To me, it’s just a convenient way to keep track of big stacks of patches. That’s enough to understand rebase and most other commands.
I recognize that this is just one example of probably many things you dislike about git, and I will not be able to change your mind about it, but:
Rebasing is essentially the user saying "actually, I wanted to make these changes on top of this different base". This is useful if you perhaps started some work on the wrong branch, or just want to test your change on top of the most recent upstream code.
> Nothing makes sense.
I think there are two key realizations to understanding git:
1. It's a directed acyclic graph. That's a very nice data structure to think and reason about.
2. Everyone can have their own version of this graph, so committing a change and sharing it with others are two different operations (and fetching others' changes is yet another).
With those things in mind:
* git rebase copies an arm of the graph and grafts it onto a different starting node
* git cherry-pick copies an arbitrary node in the graph
* git push shares your local additions to the graph with others (usually through a server)
* git fetch updates your local graph (usually from a server)
* git merge joins two arms of the graph, so that one can enjoy the changes from both at the same time
It doesn't sound as simple that way, but it's important to keep in mind that git does not track changes and instead rederives them continually.
I see git as a reluctant marriage between two disparate halves: the CLI and the ODS (on-disk structure). There might be technical reasons (and baggage) behind why the ODS is the way it is, but I don't see why the CLI is the way it is.
Outside of the time I spent, some years back, writing code to verify the porting of some mercurial repositories to git, I have no desire to look at the internals of git. Last year, I wrote a tool to query my backups stored by a software that adopted some of its ideas from git,[1] and it was not a pleasant experience. Coincidentally, that software too is just as bad at user interfaces as git is.
Why people bother getting a passport, going to an airport, going through security and sharing the ride squeezed along with hundreds of other passengers? Because the alternatives suck in comparison. My last job was one of those corporations that refused to use git because "our devs don't like CLI and git is not as intuitive as Team Foundation". Thankfully after having prospect employees and consultants laughed in their face during interviews they got the message and they switched. But those two years I had to work without git were very depressing.
Here are the reasons git was very quickly adopted:
* Better history aka not everything is a linear list of commits. Merging used to be only comparing naively two file filesystems. There was no notion of checking if changes were already applied. Now merging is reduced to the lowest amount of conflict possible; assuming you know are working along the hash enforced DAG. It was so hard that there used to people called integrators solely dedicated to merge the work of others; they still exist in some projects such as the Linux kernel but they are part of the approval process, not merging itself.
* Centralization sucked and the very notion of distributed version control didn't even exist.
* Branches are cheap. Remember, at the time there used to be flame wars about branches. This really made prototyping hard and inefficient and the only way was to copy the whole project.
* Speed. Centralized systems are inherently slow because the fact that everything is done relative to a server means that nearly every operation involve an online communication and processing from a remote machine. Git is instantaneous for 99% of the commands entered. Being fast also helps developers to not take shortcuts and making sure that what they push is clean and well divided in small and intelligible commits.
* Extreme redundancy. Having the full historic means not having to worry about "the server" getting hacked or the backups failing.
* Flexibility. How do I stash some temporary changes? How do I edit my history before pushing? How do I commit only a few lines at a time? How to I apply the changes of a few selected commits? How do I apply and check a patch file?
* Simple data structures and abstractions. There is no incompatible database schema involved here. Content is addressed and verified by hashes who reference other hashes. A 2005 git repository is still readable in 2021 and beyond. New tools and framework can and has been built around git for that very reason. Other CVS have to do their own mediocre versions.
Mercurial too has its genesis in the BitKeeper fiasco and saw an initial release the same month as git. GNU Bazaar saw a release the previous month. Fossil was released the next year.
That is what made git successful.
The git cli is, comparatively, a lot easier to grok than it was at the start.
One thing about software, once you've had a (bad) first impression it is very hard to reformulate your opinion.
If the grandparent is lucky enough to find an advocate for git, who is willing to spend time with them, perhaps their opinion might change.
That's my biggest pet peeve on this list. Git branches are really just pointers to the HEAD of some commit history, but if you've had a bunch of merges in the past it can be very difficult to tell which commit belonged to a particular branch at some point in the past.
Is there a good solution for this? Seems like it would be trivial to add a piece of metadata about which branch someone was on when a commit was made.
1. garbage collected
2. local onlyI have to admit I've never understood the point of keeping the history linear.
That wouldn’t make the complaint invalid though.
A True Scotsmans non-history-destroying version control system would reveal and be able to replay every single editing keystroke that went into a change.
If your workflow is single branch all PRs merge into main then of course you want to rebase since the only thing that matters to your history is "when did this batch of changes land in main"
But that's all local stuff. It doesn't matter if you delete something nobody else ever saw.
Branches are very often passed around to other people, examined and validated in that form. Rebasing a public branch loses some amount of important history. And even for private branches, if you rebase multiple commits at once then the testing you did between those commits gets invalidated.
Fossil allows you (for example) to fix typos in the commit messages of published branches, safely and in a way that does not destroy history and propagates cleanly with "sync". It allows you to modify the DAG safely, and in a way that does not destroy history and propagates cleanly. It does this by providing special tags that when added to a commit change the check-in comment, or branch name, or parents of the commit. The original immutable check-in is preserved, but for display purposes, the tags can override some properties of the original check-in.
Real example from just this morning: Last last night I mistakenly merged the wrong way. I merged the reuse-schema branch into trunk, rather than merging trunk into the reuse-schema branch. When I saw the problem this morning, I was able to fix it, even though this mistake had already propagated to other repositories. The change entered the DAG as a supplemental "correction" tag, so no history was lost, and if in 20 years somebody wants to go back and figure out what happened there, they can, because all information is preserved. But for day-to-day viewing of history, it looks as if I had done the merge correctly to begin with.
And I really am curious: why preserve the incorrect merge? What scenarios involve needing that information, why is that ever a good thing? If the default view of history doesn’t present these preserved mistakes, how is that really different from a rebase?
I’m also troubled by the language about “immutable” history and “destroying” history. Git doesn’t destroy any history. When you rebase, it creates a new branch. The old one is still there. You can get it back. The only reason it goes away is because it gets garbage collected later, because the user chose not to have anything pointing to it.
When I was in high school, I was taught that if I worked as a bookkeeper and I make a mistake, I should never erase the mistake. Instead, draw a line through the mistake, notate what is wrong, and enter a correction. To erase an entry in the financial ledger of a company is fraud. It is a felony. Making a correction is fine. But do not erase. Always preserve an audit trail.
I believe that VCSes should be treated similarly. While you are assembling a change, you can make as many erasures and corrections as you like. But once you commit the transaction - once you check-in the change - it then becomes part of the permanent record. To alter that transaction after the fact is akin to felony fraud. Sure, mistakes happen. By all means, correct the mistakes. But the original mistake and the correction should all be part of the audit history.
If you want to say that commits to your private branches are not part of the permanent record, and that you should therefore be permitted to edit those private branches, then I think you have a stronger case. That does not come up as much in Fossil. Fossil does support private branches, but they are seldom used. The usual case in Fossil is that all check-ins auto-sync up to the parent repo.
Shunning is not quite the same. Shunning is a mechanism for removing illegal are illegitimate content. Shunning is sometimes required to comply with legal mandates. But it is not a part of day-to-day practice. Shunning is an exception - and escape valve - undertaken only in an emergency.
Most organizations' repositories will have one or more protected branches (e.g. master). What is published there remain. Even mistakes. When the history on those branches are erased it is for very good reasons only. Usually it involves the size of the repository getting too big, illegal / private content and paths only differing by lower and upper case messing with Git on Windows. Even in those very rare cases the history of the vast majority of the files remain. And this is also something all CVS has to face so this is not specific to Git itself.
Dev branches are pushed, overwritten and erased. I don't see how that's a problem. In the end having small intelligible commits that reads like a dish recipe accelerates code reviews and corrections and I haven't seen any other VCSes doing that as efficiently as Git.
If you're using git, everyone has a local repo, and that repo's master branch is local: they can rewrite their unpublished content in it however they want. It's really quite convenient; a good feature of git.
If you commits the result of a merge before doing anything with it to fix it up, that implies that you will sometimes be committing non-buildable code which still contains conflict markers. That's not even allowed under a continuous integration policy that every build has to build (and pass unit tests and whatnot).
Git doesn’t erase mistakes. It presents a second, separate graph of commits after rebase than before. Both graphs are still there, nothing is erased, and nothing is destroyed. It’s quite important, as a VCS, that nothing is destroyed because it means if I do it wrong, I can undo.
> To erase an entry in the financial ledger of a company is fraud. It is a felony. […] I believe that VCSes should be treated similarly.
You want people who fix code mistakes to go to jail if they don’t keep a record of the mistake? Why? There’s a very, very good reason that actually lying on financial ledgers is illegal, while quietly fixing a merge mistake is not. You are conflating so many things in this broken analogy that it’s difficult to respond to.
Financial ledgers are one of the very few things in the world where history is required by law to be sacrosanct, and companies know this and agree to it in advance. The number of editable things in the world that aren’t expected to preserve history and aren’t illegal are uncountable. Nobody’s going to jail if they erase a bad chapter in a book or movie script. Nobody’s committing a felony if they tear down a house and rebuild it. Nobody is being called a deceitful liar when they erase a mistake on their math test and write down the correct answer.
Your belief was not shared by the designers of git (nor of any other DVCS before Fossil). This is the core of why your claims about git are wrong. No promise was ever made to preserve history as it happened, that was never part of the intent in its design. Therefore the very argument that git is being deceitful is a dishonest argument - it’s a projection of your personal goals onto other people and other’s people’s software, not a true story about why git was designed the way it is. The fact that your telling is motivated by trying to convince people to use Fossil over git just makes the hyperbolic framing seem extra cheesy.
> If you want to say that commits to your private branches are not part of the permanent record, and that you should therefore be permitted to edit those private branches, then I think you have a stronger case.
That is, in fact, the primary use of rebase, by a mile. I’m certain you already know this. Which is a big part of why all the hyperbole about fraud, lying, deceit is ironically not an honest narrative.
> Shunning is a mechanism for removing illegal and illegitimate content.
So is rebase. Sometimes that’s how it’s used. Maybe it happens, but I’ve never seen public history intentionally rebased other than in emergencies.
So what, exactly, is preventing people from using shun for legitimate content day-to-day? Are you policing it’s use?
Are you actually unable to engage in a significant discussion without using strawmen?
"it should be a felony" is definitely a strawman.
You are presuming to speak for @sqlite by claiming they did not mean felony, when the very example they provided of what it means to take it seriously is that it’s a crime.
What makes you certain that they weren’t suggesting exactly that?
I don’t think it’s fair to call my question a straw man. @sqlite has been making a strange moral argument out of git rebase for many years. I honestly want to understand where it’s coming from.
Fwiw I also don’t think it’s fair to ask a bullshit adhominem straw man question like “Are you actually unable to engage in a significant discussion without using strawmen?”
Stop it stop it stop it.
Stop it.
If it the program is open source, you're legally entitled to have an FTP site where only the current version of the files is available, and yesterday's version is gone. Moreover, you actually don't have to have yesterday's version retained anywhere else. Obviously, that will cause problems for you.
In between that extreme, and the other extreme (every little thing being fastidiously recorded and immutable) are rationally workable alternatives that allow reasonable mutation.
As an employed developer or contracting consultant, you may be required to adhere to certain version control practices as part of your contract. If that doesn't rule out revising history, you can revise history.
This is probably one of my biggest problems with fossil. It is very opinionated about the workflow you use. My typical workflow involves making many commits on a private branch while working on something, many of which won't be able to build or pass tests, and then clean everything up before pushing to a publicly visible branch. This allows me to easily roll back if an approach ends up not working out, or cherry-pick smaller changes to other branches if a change ends up being needed in multiple feature branches in progress at once, etc. With git that sort of workflow is trivial, with fossil, it might be possible, but it doesn't fit with fossil's blessed workflow.
Git on the otherhand is pretty flexible, and can be used for a lot of different workflows, and doesn't push you towards a specific one.
The context & argument you’re repeating here but either missing the true essence of, or teasing about, is the view that since git rebase edits commit history it is bad. The Fossil devs have long been advocating this view using hyperbolic language like saying git is ‘lying’. This unfortunately incorrect framing stems from Fossil having different goals and assumptions than git, and the Fossil devs projecting their dogmatic views on their perceived competition.
But… even the Fossil devs are not actually advocating capturing every single keystroke.
The point is people talk about rebasing and squashing a commit with a typo, and then people argue against this because it's "rewriting the history". But, as you say, why would anyone ever want to capture that typo?
He did because AFAIK he's an extremely nice guy who gladly suffers fools; I guess that's for ideological reasons, and that's his choice.
I found your own participation in the whole ordeal deeply insulting. But, again, that's me exclusively. I'm just a happy Fossil user.
Look, I think it’s absolutely great if you like Fossil. And I also think it’s great if Fossil has immutable history and cares about it deeply. My beef here is with saying that git rebase is a “lie”, being “deceitful”, “fabricating” things, etc. That is an untrue claim, it’s ironically itself completely dishonest.
If you like Fossil and want others to use it too, it might be worth reflecting on how Fossil’s marketing pages are actually pushing some people away from trying it. That’s what the claims about git really are: marketing.
The dogma that allows someone to go around calling others a liar over the tools they use to write code is what’s deeply insulting, and Dr. Hipp has had far, far more influence already than I ever will. You can stand next to this dogma if you like, but I think you’ll be on the wrong side of history. For now, blind followers everywhere like to parrot the claim that git rebase is a lie, without being equipped to defend this idea because they don’t actually share it deeply and haven’t been thoughtful enough on their own to understand the implications.
Thus why would anyone ever publish bad versions of private commits in an anonymous, unpublished branch that were locally rewritten N times?
Meanwhile, you project your disdain towards Fossil in your own hyperbolic comments and the use of intentionally divisive words like "perceived competition". Fossil is not competing with anyone.
> Fossil is a distributed version control system (DVCS) written beginning in 2007 by the architect of SQLite for the purpose of managing the SQLite project. Though Fossil was originally written specifically to support SQLite, it is now also used by countless other projects. (Fossil docs)
I'm not speaking for the Fossil project, but I'd bet that the only undesired consequence of Git's existence in their minds is the unfortunate phenomenon of ocasional Githeads stumping into Fossil's forums, foaming in the mouth about how Fossil is All That's Wrong in the Universe and how its existence is unbearable unless they change it to be more like Git.
Yeah, I can use hyperbole too. How nice!
False. This page is proof. https://www.fossil-scm.org/home/doc/trunk/www/fossil-v-git.w...
> I’d bet that the only undesired consequence of Git’s existence in their minds is the unfortunate phenomenon of ocasional [sic] Githeads stumping into Fossil’s forums
Wrong again, you lost that bet. You owe me a beer. https://www.fossil-scm.org/home/doc/trunk/www/rebaseharm.md You need to actually read the Fossil pages first to understand what this argument is even about. You are providing evidence that you’re just attacking me without understanding the history.
> I can use hyperbole too.
Do you realize that you’re totally validating my whole argument? Because you’re calling my argument hyperbole and a straw man, when I’m accurately describing what Fossil and it’s author have said about git rebase, without realizing it you are agreeing with me and disagreeing with @SQLite. You claimed that my question about rebase being a felony was a straw man, when in fact that’s exactly what @SQLite said literally. You called @SQLite’s idea ridiculous, not mine. You are disagreeing with the core belief that Dr. Hipp holds that modifying code history should be at the level of a criminal offense.
I’m upvoting you because you are proving my point.
I don't see how that helps the problem.
One of git's goals is to record the history of a project, and one can argue that that feature can be improved.
I think recording every typo would be worse, as it becomes harder to find the edit you're looking for. More hay covering the same needle.
I don't think rebasing is a bad feature, but I think merges and branches in the history provide o more useful story of how the code was written.
If a branch is public and is merged, then we need to preserve the original and track the relationship.
Git is fundamentally broken with both rebase and merge.
If we consider a branch which has a single commit, then rebase and merge achieve something very similar. The text is merged and a new commit is created. The target branch head is updated to point to that commit, which has the previous head as its leftmost parent.
The merge version will record two parents, whereas the rebase (actually a cherry-pick under the hood) records only the destination branch parent. This omission is unfortunate for the rebase.
What people do when rebasing an already published commit is put something in the commit message like "Cherry-picked from <SHA>", which really wants to be a parent pointer in the meta-data, so that it shows up under "git log --graph" and such.
On the other hand, if we merge multiple commits, rather than rebase, another idiotic thing happens: a single, squashed commit gets created out of the merge result and planted into the target branch.
We need a flavor of git rebase which records a parent pointer for every single cherry-picked commit. That will create a merge which makes sense. Merge a feature branch with 7 commits, and you get 7 individually merged commits, each with a parent. (Or, maybe you will just get six: perhaps one of them became obsolete and was skipped.)
Also, if commits have additional parent pointers, then any editing workflows involving rebase should preserve them. Say I locally merged the foo-feature branch into master. I now have 7 new unpublished commits coming from the merge. But, uh oh, the merge was actually bad. While textually it went fine, something is wrong and needs to be fixed up. Say I need to reach for interactive rebase to manipulate these 7 commits; maybe adjust something in the third one. This interactive rebase job should preserve the above-described parentage.
BTW, Fossil has commands called “shun” and “scrub” that do effectively rewrite history. Every version control system has them, because you have to have them. Being on a high horse over rebase is 1- not going to change rebase, and 2- makes it seem like you don’t fully understand git, and it’s design decisions and goals, therefore not qualified to opine on it. See Chesterton’s Fence.
All of this is mostly irrelevant to getting coding done. Wanting immutable history is fine, if that’s what you want. I’ve never heard a compelling reason to need it, and it’s not one of git’s goals, and never was. (Nor was it a goal of any other version control system either, not of svn or p4 or hg or source safe…) Git is version control, and it does what it was designed to do which is keep multiple versions around, let you share versions with other people, and be a safety net against code loss. I can’t think of a single time in my decades of professional software career where it mattered that someone edited their local history before pushing, and/or didn’t preserve some strange sacrosanct history of the exact order every line of code was written. I don’t have anything against holding this goal, but I’m unconvinced that it’s relevant.
> I don’t have anything against holding this goal, but I’m unconvinced that it’s relevant.
I am likewise not convinced of the value of immutability.
I see the git repository as an object being edited, and that object has no history of its own. When we add a new commit, say, we are making a new version of the whole git-object which now has some new objects in it, plus a moved head.
The old git repo in which that head pointed to the previous commit doesn't exist any more. Well, some people have copies of it. They can stick with those, or get this one. Take it or leave it.
We should be able to control any aspect of our work. If I want to rewrite thousands of commit messages in the published master branch, that's my privilege. Git deals with this just fine; someone who fetches that is informed that their master is divergent by thousands of commits. They can disagree with what I did and then just cherry pick things onto the original branch going forward.
Bisect, blame, and log, among other features, are all easier to use with a linear history. They’re not unusable with non-linear / many branches, just easier linear.
I’m not quite sure I understand your comment though - it implies that you consider linear history to be a strange and/or non-default thing to do. What’s the implied point of not keeping the history linear, or what is your preferred workflow? I’m not entirely sure what’s being compared, but linear is the default you get from more than 1 commit. Linear is very natural with a team of one or only a very small number of people. Linear is a lot nicer than branchy for branches that have exactly 1 commit.
> I’m not quite sure I understand your comment though - it implies that you consider linear history to be a strange and/or non-default thing to do.
Sometimes you get a branch, I get them even on personal projects if I forget to push and work on a different computer.
I've heard people say that, when they do get a branch, they will rebase to make the history linear instead of having a merge. I don't think that is default.
1. Always use merge commits (--no-ff) a la git flow, and leave the branch name in the merge commit. That doesn't get you the branch name on the commit, but it should help identify the branches.
2. Use a pre-commit hook to put the branch name or bug tracker ID at the bottom of the commit message. There are some tools built around this, I've done it at $DAYJOB with a custom python script I manually copy around into `.git/hooks/pre-commit`.
I saw a new dev do this recently. They dug into the original summary of the changes from the dev (for a PR that landed 6 months before they joined) and the flow of conversation around the work. They had insights into the decisions the team made around the feature as it was developed. The PR linked to the product artifacts that led to the work, so they had the business context. They nailed the implementation, first time :-)
There's a nice benefit to doing merge/squash and providing a link to a set of context around the specific set of commits, not just a contextless line of commits, as clean as it may be.
What if the branch has 17 commits, the 9th of which introduces a regression? Moreover, suppose that the 9th commit of the original branch didn't exhibit the regression; rather, its change had that effect when it was merged into master.
You need all 17 commits, in their rebased versions, on master, so you can search through their history (e.g. with git bisect) and discover that the 9th one broke it.
Rebase is the correct thing: it calculates a new sequence of commits that are individually merged, and retained as separate commits that you can dig through in the new branch.
What's missing is this: the rebase of an entire branch into a new branch should record a second commit parent (for the final commit of the operation), pointing to the final commit of the original branch.
Like this:
X-Y-Z--<< branch
/
-A-B-C-D-E--<< master
rebase it: X-Y-Z--<< branch
/ \___
/ \
-A-B-C-D-E-X'-Y'-Z'---<< master
where X', Y', Z' are the rebased versions of X, Y, Z.A regular git rebase is missing that second parent link from Z' to Z. You just get this:
X-Y-Z--<< branch
/
/
-A-B-C-D-E-X'-Y'-Z'---<< master
The only way you can identify that Z' was cherry-picked from Z is by clues in the commit message: identical commit message text, or something like a Gerrit Change-Id.Possibly, each cherry-picked commit should have the original as its parent:
X-Y-Z--<< branch
/ \_\_\___
/ \ \ \
-A-B-C-D-E-X'-Y'-Z'---<< master
now that is almost like iterating over the branch and doing a commit-by-commit merge.At first we pretend that the branch is just X, and merge it to get X'. Then we merge Y', and then Z'. We never merge a sequence of two or more commits into one.
If you do that, you have both: nobody can say you're not merging, but you have the effect of a rebase.
A rebase (cherry-pick) is a kid of merge: of one commit (at a time), without recording all the parents.
A case can be made for rebasing commit-at-at-time using merge. (Not for private work, obviously; just already published, permanent history).
The nice thing about rebase is that it doesn't introduce the crufty parentage, which makes it ideal for rearranging unpublished history and then presenting a clutter-free end result.
For some developers, the commit is the cohesive unit of change. For others, the PR as a whole is the cohesive unit of change. I fall into the second group. The ninth commit of my branch is meaningless -- it's only the branch as a whole which can be understood as a single coherent thing. Squash merges reflect and support my methodology.
Yes! This is ideal.
Totally disagree, because the whole reason for this is to support multiple developers working simultaneously of the same master branch. That is, sure, master may just be a linear history of commits, but at any specific point in time there may be 10 feature branches off master, each from a different branch point.
(Gitwise-factually, a commit is a snapshot, not a change, but a new snapshot represents a change relative to some other snapshot, like any one of its indicated parents.)
Your "methodology" is not rooted in facts.
It can be described as wanting every version control to be CVS.
Back in the day when those kinds of systems were widely used, the most common complaint from knowledgeable CM people was that there are no change sets: related sets of changes treated as a unit which contains the fine-grained changes, with individual descriptions, references to tickets and so on.
If we regard a feature branch as a change set, when you do a squashing merge in git, you're vandalizing the change set; the merge you're doing is only slightly better than "cvs -j from -j to" (merge the delta between from and to onto the working copy). How it's slightly better is that it records the extra parents, for tracking.
One of the main improvements in git over things like CVS and Subversion is that you can at least halfway simulate change sets, due to the relative ease of manipulating individual commits and the scripting done around it (like git rebase scripted around cherry-pick).
> If we regard a feature branch as a change set, when you do a squashing merge in git, you're vandalizing the change set
I'm not. The change set is atomic, the individual commits are essentially implementation details which the PR encapsulates, and when it's merged to main those details become opaque. I mean this is just my methodology, it doesn't have to work for everyone but it's the natural expression of my work for me.
Be that as it may, a squashing merge is focused on one level of abstraction and mostly poops over the others.
The nice thing about levels of abstraction is that we can have them all and preserve them.
Commits which you did not actually produce as part of a merge are not an implementation detail; something which isn't there does not provide implementation. Opaque implementation details are something which exists; they can be looked at debugged into.
Change-Id: I7321a7fddab667af7a418eeeb7d48c0240570260
The Change-Id will be the same whenever that commit is cherry-picked anywhere, even if it has to change due to merging. It is not a hash, but a random UUID.If you trowel through the log messages you can find the places where that change has "been" by looking for the Id.
The idea is ... interesting, but the implementation of Change-Id has quite a few problematic corner-cases which make normal git operations (rebase, cherry-pick, merge) more delicate than the need to be.
I often make commits without being on a branch at all (detached head). I like this about Git.
that's it, right there. Why are you bothering me about a detached head?
Yeah, it would make you `git branch` to give it a name, or an override option to commit it without a name
GIT - the stupid content tracker
"git" can mean anything, depending on your mood.
- random three-letter combination that is pronounceable, and not
actually used by any common UNIX command. The fact that it is a
mispronunciation of "get" may or may not be relevant.
- stupid. contemptible and despicable. simple. Take your pick from the
dictionary of slang.
- "global information tracker": you're in a good mood, and it actually
works for you. Angels sing, and a light suddenly fills the room.
- "goddamn idiotic truckload of sh*t": when it breaks
This is a stupid (but extremely fast) directory content manager. It
doesn't do a whole lot, but what it _does_ do is track directory
contents efficiently.
- README, Linus Torvalds, Git source code, April 7th 2005It's not a "passion project" by any means. It's a highly developed, well engineered piece of software and there's lots of other non "passion projects" that would do well to learn from it.
I don't mean to disparage SQLite in any way, it is quite the achievement, but your characterization is mistaken.
The fundamental difference is: sqlite is a cathedral, git is for bazaars. That's kind of a big deal.
Why would you use git if you could use Mercurial instead?
GitHub would have won even if it were MercurialHub. It had a better user model of the problem to be solved than any other SaaS code forge, and it supported both open-source bazaar-style development and private enterprise. The gamification and social network features are KILLER.
Also, the Mercurial developers never had the built-in fan base of Linus Torvalds. Anyone who wanted to work on Linux had to use Git, so that was another major reason to learn (and stick with) Git.
This. Which basically means: no time to setup, no operations overhead, no administration required. This means, no dedicated team member needed who is the "admin" for version control, no dedicated infra required, etc. In the early 2000s when we were a team of 8 developers, one was a "ClearCase admin". He was a single point of contact for all things ClearCase (request him to create a new branch, add a new user, create a label, create a new repository and so on). He also had to ensure regular backups were done, the server was regularly patched, the disk had enough space. GitHub/Gitlab simply remove all of this just for a few dollars per user.
- But with mercurial branching where more like forks with different directories. Made me avoid that feature in web development work (that way I didn't have to fiddle with my webserver config every time I would have liked to create a branch)
- I recall having to deal with a bunch of gotchas around mercurial bookmarks (tags), like not being able to delete them, problems with sharing them with others, and things like that.
I was actively exploring different DVCS around that time, coming from a SVN background where I remember spending a lot of time fixing merge conflicts like a sucker. The one I liked, for its theoretical aspects, was darcs (as I was also leaning heavily in the Haskell ecosystem of the time).
Anyway, in short for me git overtook mercurial because it was easy for me to hop around multiple branches, pick off individual commits, and rebase commits (I was still pedantic about how I formatted my commits before a PR, nowadays, not so much).
That's bzr, not Mercurial. Mercurial has always had more advanced branches than Git.
I'm sure a lot has changed in that timespan, and maybe it was a bit more of a nuanced situation wrt branching, but the pain points are the things I can only vaguely recall at this point.
Because Mercurial was (Still is?) hopelessly slow on large repositories.
Mercurial did have one killer feature, and that was hginit.com. Git's documentation was terrible for a long time. I would send programmers used to old fashion version control to hginit all the time, even when I was teaching Git.
- The fact that it was Linus developing it
- The fact that it was released just after Git
- It's pretty much like Git and evolved to be usable like Git
- Git had way more advanced tooling and still does to this day
Why SQLite Does Not Use Git - https://news.ycombinator.com/item?id=16806114 - April 2018 (608 comments)
also this bit:
Why SQLite Does Not Use Git? - https://news.ycombinator.com/item?id=26112802 - Feb 2021 (1 comment)
News flash: Many programmers love complexity.
Complexity == popularity == profit
Prove me wrong, please (I seek simpler stuff)
I used to get into exploring every nook and cranny of every config file and endlessly tweaking settings.
Now I strive to use everything with default settings if at all possible.
Not sure if you're trying to argue that Linus was a "new/inexperienced" programmer when he first wrote git.
I think the "it's too complex" argument against git is misplaced, or at least not complete. I don't find git "complicated for complexity's" sake (e.g. looking at Enterprise Java Beans back in the day), but it has a very specific workflow in mind for its conceptual model, and if you don't understand that conceptual model, you're just going to be confused.
It took me a while to grok that model, but once I did it clicked and I understood it. Not saying it's optimal, but it also has a lot of benefits.
This is an 18wheeler that can haul anything but for "basic" dev doing simple administrative apps, a bicycle VCS is more than enough. A lot of dev I've met don't even understand the deep ramifications of multiples branches with any VCS so understanding the complexity of git is way above what we require of them.
Git has a lot of benefits but I'm not quite sure that it's really a good thing that it is slowly becoming the de facto standard.
Wait and see...
Now I want my operating system and my tools to shut up and get out of my way because whatever irrelevant thing they think is important is not what I ever need to be working on.
Everyone will agree with us of course, but suggest something minor like "hey maybe you don't need react" and here come the reservations. To paraphrase Alan Kay, everybody likes simplicity except for the simplicity part.
It is easier to design complex systems and harder to simplify. Also complex systems have positive economic incentives like consulting training books or that if simple will not be sustainable.
"Git doesn't do X so I'll just make this tool to add the feature"
One hundred tools and a few consolidations into toolkits later you have a massively complex tool chain scope creeped from some much smaller problem deep down in git
Loving complexity for the sake of complexity or sake of job-security or competitive advantage covers only small subset of situations.
There’re a lot more cases where seemingly simple problem has a lot of hidden complexity, and only complex enough solutions succeed at successfully solving it and staying competitive on the market…
We don't use DVCSes for their own sake, we use them to accomplish something else. To the extent that the DVCS gets in the way, it's an opportunity cost.
Entire node/js/ts ecosystem resonates strongly here. They reinvent the wheel, assign it a new name, thought-leaders market those inventions to every corner of the tech ecosystem, and then they add another few layers to their build chain.
Maybe that makes me lazy (which I am) or a bad developer. It is a barrier though as, like it or not, we all know git.
This, in a nutshell, sums up my feelings about git.
CVS and Visual Source Safe way back in the day.
I use it nowadays, but I think it's too limited to provide value like Perforce changelists does.
I'd like to have one changelist for random local dev changes, like setting debug flags or turning off optimizations. Stuff I have no plan on checking in.
Then I'd like to group my edits into maybe one to three possible future commits.
The staging area doesn't allow this, I need to keep doing "git add -p" and carefully sidestep all the debugging stuff, plus that I can only work on one commit at a time.
So could have been useful, but not so useful in it's current state IMO.
This is that part of git that bugs me the most! It's totally hostile to the way I prefer to work: one screen session containing all the editors and shells per change that I am concurrently working on. Git simply cannot do it, unless I clone the entire repo for every change, which is not practical for large projects.
Note that you can have multiple working copies from one clone for a few years now. git worktree add <path> <branch>
# At the point I realize I have two commits
git add -p
git commit -m 'first'
git commit -am 'second'
# Do some more editing
git add -p; git commit -m 'debug' # commit the debug stuff so that I don't have to keep skipping it
git add -p; git commit -m 'first'
git add -p; git commit -m 'second'
# After repeating the above a few times
git rebase -i origin/master
# Remove the debug commits, move the first/second/etc. commits together, squash the runs of first/second/etc. into one commit each, and then go back and write commit messages for each of themI found darcs to be the most ergonomic, except for the fact that it was orders of magnitude too slow.
SVN was good in that it fixed most of the pain-points that CVS had.
git and hg are both great. I tried them both out at about the same time and found hg to be more intuitive, but it didn't take me long using git to be comfortable with it.
The UI of git is honestly quite terrible. Commands are often confusingly named and have seemingly disparate uses. However the underlying model of git is straightforward enough for me to reason about what operations should be possible.
Pretty much all VCSs require some such knowledge and reasoning; e.g.:
- The lack of atomic commits in CVS stems from the underlying RCS stack it is based on.
- SVNs brain-dead branching stems from the fact that there are no branches in SVN, just O(1) copies from one path to another.
I know/use clone, diff, add, commit, and pull/push.
That is all I ever do with Git, and it's pretty straightforward. I have found no need for the rest of it, and it confused me when I tried to look into it.
Note: I use git frequently every day, not by choice, but simply because it is the defacto version tracking system. I have grown to like it, but I miss the more straightforward every day use of Mercurial (which is rarely supported by any services anymore).
Edit: ... and as a native English speaker, I can't help but feel bad for others who are expected to become familiar with these functions without familiarity with these colloquialisms.
But the command line user interface is one of the worst I've ever seen. The names are bad, discoverability is poor and which command line tool does what doesn't always make sense.
The requirements of Linus and the kernel developers are completely different from the requirements of the vast majority of users of Git, who are using it in small teams in a centralized way.
I would argue that the only benefit that this decentralization gives is a local cache, which makes things nice and fast. Most users don't need to be able to create branches easily (code review tools already let you do that, effectively) and it's just given people enough rope to hang themselves (e.g. GitFlow).
IMO the worst aspect of this needless complexity is having to explain to people "you've got 1) the branch central repo, 2) your local copy of that branch, and 3) your local branch". It gets even more complicated when personal forks come into the picture, as there are just so many copies of things to manage or trip over.
Try explaining to a junior developer why "git fetch" takes two arguments ("origin" and "master") while "git rebase" takes just one: "origin/master".
It's sort of how a lot of shops decide they need to incorporate 'AI' into everything, or they need to use microservices and Kubernetes to solve their problem. Some of it is just resume driven development but a lot of it is just observing and following trends.
> Some of it is just resume driven development but a lot of it is just observing and following trends.
Couldn't agree more. That's exactly what I think happened with git.
The command line is shit, one example: `git checkout` can both switch to a new branch or restore a file.
There are certainly cases where Git is the right tool, but it's vastly overapplied, used because it's popular, not because it's the best tool for every job. For smaller projects without a lot of randos offering PRs from the outside, Git's probably costing you more in its complexity than it's saving you by having that complexity at hand.
One of Fossil's selling points is that it's missing several important bits (i.e. all the tools for rewriting history)
The issue is that everyone is using the most complex tool, from solo developer, to 10 developer teams, to 1000 developer projects. You should use the simplest tool that gets the job done for your project[0]. That isn't git, but it's so easy to just pick the thing that "scales".
[0]: https://blog.codinghorror.com/the-principle-of-least-power/
Cherry pick on the other hand has a non-technical definition that closely approximates what it does in git: https://www.merriam-webster.com/dictionary/cherry-pick
EDIT: cherry picking is actually making a copy of the leaf and gluing it elsewhere. And fast forwarding isn’t really moving the leaf, it’s taking the leaf you call BRANCHNAME and now calling a leaf further up the branch BRANCHNAME.
https://www.git-scm.com/book/en/v2/Git-Tools-Rerere
> Plus that it's quite hard to rewrite dates, names, or split commits into several.
Read the documentation for 'rebase' and 'cherry-pick' and 'add'.
> https://www.git-scm.com/book/en/v2/Git-Tools-Rerere
Thank you for this, looks like exactly what I need at the moment maintaining a dev branch alongside a bugfix deployed main branch and having to do lots of rebasing.
Once users figured out what they want to do with the graphs, they then need to figure out how to achieve that with git's cli. And I feel that is where many of the complaints about git's complexities comes from.
IMHO the best way to use git is through a GUI so that one only needs to figure out what they want to do with the commit graphs, leaving the second part of the puzzle to the GUI.
Git set out to solve a different problem: adapt VCS to an agile, CI/CD world. It does this much better than those previous VCS generations, but its mental model does seem needlessly complex.
In order to achieve that, I think you end up with some additional complexity like local vs remote and a complicated graph structure to accommodate people coming "on" and "off" line
Yeah, I guess I kind of consider that an agile tenet: in general, that people can work quickly and iteratively with limited interdependencies.
Said another way, it would be difficult for teams to work in an agile fashion--where the work itself is frequently distributed and happening in parallel--using most older VCS models.
Main point being that comparing the complexity of Git and, say, SVN isn't apples to apples, as the former is suited to use cases that are inherently more complex.
Which were "send patches to the maintainer", and a little bit of "work offline".
Linus fixed those as needed, and people who haven't those problems are using his solution a bit like painful fashionable shoes - they hurt but dagnabbit, that's what the fashion is ...
> Git’s internal model is trivially simple: everything is stored as objects addressed by content. These objects can point to each other. Additionally names can be associated with objects and which object a name points to can be updated.
I'll do you one better: it's just some byte arrays on disk.
I think the OP has a section that gets at some of the complexity of git's model:
> The complexity of Git distracts attention from the software under development. A user of Git needs to keep all of the following in mind:
> a. The working directory
> b. The "index" or staging area
> c. The local head
> d. The local copy of the remote head
> e. The actual remote head
> Git has commands (and/or options on commands) for moving and comparing content between all of these locations.
Maybe they fixed that in the 20 years since i used it.. but that was a disaster.
That said, Mercurial and Fossil came out within a year of GIT so GIT at least outcompeted them early on.
Git, Mercurial, and Bazaar all came out within three weeks of each other (late March/early April 2005).
For example, why can't I do something like
git view filename -3 # shows file before last 3 changes _made_to_the_file_ (irrespective of how many commits)
git view filename $date # shows state of file at date
git revert filename $date | number # roll back a specific file to what it looked like at $date or file version
I have the feeling that the developers of git even understand this, but they needed to get git out the door and assumed the community would build the layer atop git to make it more user-friendly. Sadly, attempts to make git user-friendly have so far boiled down to GUIs rather than introducing new concepts. Certainly the current collection of shell scripts that exist as the ui for git definately have a hackish feel to them.… of what, though? On GitHub, and on every other project I've worked on that used git, the natural unit is a set of changes (e.g. a GitHub PR, a patch). If you run `git show $SHA`, git will show you a set of changes. But that's a lie; that's not what git stores.
If you run `git log`, it'll give you a list of commits. If you run `git show` for each, it'll show you a bunch of changes. Many users expect (quite reasonably, before they've been bitten) that if you start from scratch and apply those changes in order, what you end up with will be the same as your current tree. Nope.
What "natural" does even mean? That's a vague philosophical argument. And even then, all versioning systems deal with versions, not changes. Changes are relative to different versions. This is true even for SVN. Changes are always relative, it doesn't make sense to track them as the based unit conceptually or even for performance. Different versions implies change, not the other way around.
> If you run `git show $SHA`, git will show you a set of changes. But that's a lie; that's not what git stores.
Each new version will bring a change. This is not a lie. Sure it could show the entire corresponding tree by flooding your console, or it could be more practical and merely show the diff from the previous parent(s). This is the kind of thing that people who do not bother to take the time (about one hour) are stuck with. Don't worry, it's not that hard.
> Many users expect (quite reasonably, before they've been bitten) that if you start from scratch and apply those changes in order, what you end up with will be the same as your current tree. Nope.
The author, the date and other data about the commit do matter. If you rebase your commits, which is effectively applying those changes in order, then very very likely you will not have the same snapshots of the filesystem. And even if the filesystem remain the same, the data about the commit will be changed. It seems that your expectations of what a CVS "truly stores" is naive and based on a lack of experience.
No, Darcs and Pijul (for which I have high hopes, though it's not ready for prime time) deal with changes.
> It seems that your expectations of what a CVS "truly stores" is naive and based on a lack of experience.
Thanks, but attempted insults notwithstanding, I know what CVS truly stores, and what SVN, RCS, SCCS, git, Mercurial, darcs, and pijul truly store. (I admit I don't know what Perforce or ClearCase truly store, or others I may have used too little to recall.) And before SCCS I knew what a stack of labelled backup tapes truly stored, though I wasn't _quite_ as proud as git of my fancy labels.
I didn't know about them. From what I read, Darcks is slow and Pijul doesn't really bring anything new to git and seems to "solve" the problem not being based on patches while not really doing better at what git already solved. At the end of the day, both can compare what was applied and what wasn't but the algorithms used around it are just more complicated no good reason, this means less tools and less community support.
Then, Darcs has never been slow for me, I've been using it for about 15 years, I can remember of one minor performance issue. It doesn't scale well to very large projects, but there are extremely few of them. Since concrete arguments are important when discussing with others, the main event which started Darcs' decline was its inability to handle conflicts in GHC (the Haskell compiler), which was at the time the open source project with the largest history (as measured in number of patches).
As for Pijul, I'm curious where you got your information from. I know Pijul's authors well, at least enough to know that they are well aware of the algorithms used in Git. Pijul does solve a number of problems:
- I know at least one very large-scale project where collaborators are well aware that 3-way merge can mess with their code in really bad ways, and complain about wasting tons of time with that. Unfortunately, I can't embarrass the project's authors by sharing their name, but I can say that it is a widely used piece of software with a really long history.
- Knowing that the code you review is the code you merge is the only thing you should care about if you care about your code at all. This is not what 3-way merge gives you, as the Pijul manual clearly explains (see "lack of associativity").
- Patch commutation is what Git tries to do all the time (rebase, merge, rerere…), except with bad hacks which fail in the more complex situations, which is where we need them the most.
Then, I don't think you actually understand what we mean by patches: for anyone who has looked into the topic, patches are obviously not equivalent to commits except in the most basic cases (where all tools "work", in your own terms).
A datastructure is defined only with respect to the operations you define on it, I'm sure you use other commands than checkout. For example, people who review snapshots rather than patches before a merge (or even a commit!) either have really small projects, or way too many man-hours to spend.
> if you take one or two hours to learn the documentation is actually intuitive
Git isn't "intuitive" at all for anything but the most basic operations (I agree it is easy for those, if you are willing to work with snapshots). For example, the existence of `git rerere` clearly shows that Git doesn't model conflicts at all. Then, merges are not unique, Git chooses one at random when there are more than one. Rebase is a horrible hack, "replaying" commits doesn't even always work.
Moreover, saying Git is easy isn't a very effective way to show off on HN: if you feel you have to defend Git so strongly, even though it wasn't actually attacked, and to the point of refusing to reply directly to any rational argument, it clearly shows you invested enough time in Git to feel you would lose something if you had to move to a better solution. So, sorry, but I don't think it was just "one or two hours" for you.
> git gets easier once you get the basic idea that branches are homeomorphic endofunctors mapping submanifolds of a Hilbert space.
Mercurial names are much more natural but it lost the war already.
I use git every day and don't complain when using it. But when a junior is learning git, you start seeing the oddities.
It's starts by saying that it's not comparing git and fossil but just keeps trying to throw generalities at your face how fossil is so much better and smarter than git.
It sometimes compares git and all of a sudden compares github or gitlab, no reason given. So what is it about, comparing git, github or gitlab?
Apparently git is to complex for developers, sorry but I have yet to see a developer not managing to work with git. Either I've always worked with geniuses or this is a non-problem.
There's options like gitea that are lightweight and super easy to host.
You must work with geniuses.
No one I know has never been at the place where you just accept the tree is fucked, delete .git and paste your changes back into a fresh checkout.
Granted now I understand things like git reflog so it's been many years since then, but git is as far from approachable as tooling gets...
I've never seen a tree being broken, and needing this kind of action.
Don't get me wrong, the git cli could definitely be improved, but as long as you have a clear workflow and map it to git actions it works well.
This is a local issue that usually bites a new developer.
-
Sometimes you can search/ask for a fix, sometimes you realize in the time it took to find that fix you could have just made a new local copy.
I was a "productive developer" capable of solving real issues for a pretty long time before I was comfortable with "advanced Git", mostly from writing tooling with it.
And there are plenty of brilliant developers who never really get past the basics.
It seems pretty backwards that the tooling mainly meant for the background tasks like backing up and merging changes should be harder to master than the foreground tasks of actually building the stuff you put in Git...
Git is not hard, most developers, from junior to experienced do fine with it with really little learning on their own, no need to micromanage anybody.
I literally said these are problems a manager will not witness unless they're micromanaging, where did I judge your management style?
-
Though to be blunt, this is starting to feel like the Dunning-Kruger effect in action. I mean how deeply familiar can you be with Git and still question how others find Git hard?
Even Linus has admitted that Git is hard, but argues it's gotten easier over time. Of course things live on the internet forever, so even fresh beginners end up not learning about the newer, easier ways to do things.
I think the mentions of software like GitHub were more in terms of workflow since apparently Fossil bundles all that together in a static binary (you don't have to integrate the tool chain yourself)
That's actually the thing that makes git hard to replace, there's a lot of tooling built around it.
I'm not sure that is really the case. Almost everything in the git universe is git + webserver + cli tool + scripts + gui tool. It really isn't the same... you have to concatenate an awful lot of tools to add fossil's base features to git. That said, I end up using git on everything but personal projects... then I use fossil.
1. Pull Requests ( https://fossil-scm.org/forum/forumpost/01d77b7259ea0b00?t=h )
2. Code Reviews -- I spend atleast half my time doing code reviews these days and the tools just aren't there for Fossil, GitHub.com has good tools
3. "git add -p" -- I try to make commits a logical change, often figuring out what I'm doing involves making a set of physical transforms and then I like to commit them as logical transforms
I think the short list of git grievances here are outweighed by the costs of not using git.
Seems like git=(gitlab|github) for the author. I don't agree this is right, as these both are totally different products comparing to git and made with different purposes and aims.
Phone access. Okay, probably we are in different culture contexts, but for me a need to use my phone to access source code (or tickets) always means things went very very wrong.
The author may have a point about having better views of the source from GH/GL, but they are absolutely providing a decent experience for it already.
It's clear that both projects need a revision control to manage changes, yet the requirements seem to diverge. Linux kernel is a massive project with a lot of involved contributors. While SQLite is a tighter team, majority of commits are reasonably by 'drh'. And the codebase itself, of course, is not on the same span scale.
As such, my understanding is that Git needed to allow slicing and dicing of various sources of patches so that they could be properly sequenced, tested, reviewed, and integrated (thus index/staging and rebase). On the other hand, SQLite needed an immutable history as part of its exhaustive approach to testing and documentation, and Fossil could also be an excellent testing ground for SQLite too ... and a more cohesive CLI (thus the repo-db, security model, persistent branch names, non-destructive delete etc.).
So, just as great Masters do, they created the right tools for their project. AND they shared their tools with the rest of the world, that is us all, for which we are grateful indeed.
Now, is Git or Fossil a right tool for your project depends on your needs and choices. Not all projects need the slicing and dicing of the patches, just as not all projects require immutability of the history. If anything, the current reality is that Git became a sort of C of the revision control, just it's possible to have a practical use of it without the need to learn it thoroughly. The requirement to rebase your changes for a PR is rather imposed by project conventions, not the tool.
I like to think about Git as being a Swiss knife to carve the wooden blocks of patches, so they fit together. Meanwhile, Fossil is a kind of a handy set of chisels to etch the revisions into the stone tablets of the project.
We get it. Many people don’t like it. They prefer to use $TOOL instead.
It’s been 16 years. Can’t we resign this to emacs vs vi category of things we’ve largely accepted and stopped arguing about yet?
I’m not mad at the article. They have preference and wrote down why they don’t use the more popular tool for when people ask. That’s fine. But I’m guessing it was posted to have the “discussion” we’re having right now (and every time got comes up). People bashing (or defending) git.
And it’s exhausting.
We’re half way to the age of vim. Can we let this one go too please?
Additionally, it is cathartic (and occasionally even productive) to complain about the warts and sharp edges of a popular tool, especially when the tool is the overwhelming the tool of choice in its domain.
To which I can only reply: No.
Yes I'm sure there's always a way to recover using the tools, but if you are not a git guru and don't have one handy, deleting and cloning is usually the more pragmatic way.
mkdir ~/fossils
fossil clone https://sqlite.org/src ~/fossils/sqlite.fossil
mkdir ~/sqlite; cd ~/sqlite
fossil open ~/fossils/sqlite.fossilOne way to measure the over-complexity of software might be:
1. How easy it is to screw up
2. How hard it is to fix things once you've screwed up.
Git doesn't score well by that test IMO.
Back when I first learned Git, I thought it needed an undo, at least for the local repository. Not a revert commit or partial/piecemeal undo this or that, but an "I screwed up, take me back to x point in time or snapshot" level undo.
Ironically, Google & Facebook choose Hg over Git because they can make Hg run much faster than Git at their scale: While both vanilla Git and Hg are too slow for their huge monorepo, they basically replaced most of the core of Hg with their own implementation to speed it up for internal use, keeping only the Hg cli.
The builtin web interface is very nice, too.
I have since had to switch to git at work, and after discovering magit I don't think I'd want to go back. But Fossil has been very pleasant to use.
One thing that’s not well appreciated even with most people who use git is how much safety net it can provide against losing changes accidentally. I know a lot of programmers who don’t care and will blow away and re-clone a repo before learning how to use git reflog. But, in git you can recover from almost anything if you know how.
P4 is pretty terrible when you want to try something multiple ways, or when you want a small team to prototype a feature without affecting everyone. I’m actively using both git and p4, and I can’t tell you how many p4 shelves I’ve got stacked up. They keep collecting because there’s nowhere else to put them, and it’s the only (crappy) way to test parallel ideas.
Using git for trying out things in different ways is so nice and so fast once you get fluent with it. Of course your stance is valid and perfectly fine, you don’t have to change away from p4, but I think you might be surprised how git changes your thinking about source control.
BTW I haven’t used Fossil or hg or plastic or other new source control systems, but I would assume you could replace any of them with git in my above text.
Git has a lot of weird stuff and fluff, but that doesn't mean you have to use all of them.
For instance, I try to almost never branch and use deployment flags instead - then you can always rebase and maintain an extremely simple commit history.
Simple >> Complex.
The biggest benefit of this approach imo is the drastically reduced chance of a bad feature going out because someone had worked on it for 5 months in isolation and it becomes virtually impossible to debug.
Also feature flags are not going to be able to cover all code changes. For example, it's hard to put code refactoring behind a feature flag.
Finally, wide-spread use of feature flags as a type of in-code branch certainly seems like code-smell to me. It's fine for a while, but after 20 years you're going to have a difficult time maintaining that repo unless it's done with sufficient discipline that those feature flags are quickly removed once the feature hits production. It will also screw with your automated static analysis tooling.
Personally, I like smaller, well defined changes and CI/CD.
This is true. But this is a criticism of Git**b not git.
"Use git checkout."
"And how should I create a branch?"
"Use git checkout."
"And how should I update the contents of a single file in my working directory, without involving branches at all?"
"Use git checkout."
It actually makes sense if you learn the full command but makes zero sense when you google individual common operations and then they thought "this command should also have a shorthand to create branches" I suppose
`git checkout` for restoring files is dumb though, but as others have mentioned the UX has been improved with `git restore`
I call this "hacker tunnel vision". It is a state afflicting a lot of programmers at the top of their game when it comes to writing end-user API and documentation. It can be summarized by "it is so obvious that people will think I'm mocking them for explaining this"
Git checkout covering multiple usecases makes sense to me, I've never thought, hey this is used for all these different things they should be separate. I think, hey, this tool is always checking out a hash and a path.
The fact that can be a branch or a commit is entirely inconsequential. If it irks you it's because you think there is a difference between a branch and a commit then you're likely lacking some fundamental understanding about git.
I'd even argue that hiding this fact is what makes all of the other CVS a pain in the ass to use because they are trying so hard to hide the truth from their users.
For example SVN is strictly a linear snapshot of a single tree.
I used svn for a long time. New developers got it right away. Git is treacherous waters for 7/10 of the developers I've worked with.
That said, I LOVE GitHub and gitlab. Amazing tools for collaboration. But the Git part doesn't really do it for me. I'm waiting for something better to take its place.
A DAG is one of the simplest data structure out there. What is complicated is doing merges in non-DAG CVS and let's not even talk about more "complex" operations like putting all of your commits on top of a branch, only committing a single line and, god forbids, creating branches without asking for permission.
> over-engineered solution to problems that should be simple
I get it that you're unfamiliar with how not nice databases schemas are. A git repository from 2005 will still work in 2021. AFAIK all other CVS will fail at this because they rely the overly complicated data representation. Git's internal representation is very simple. Filesystem snapshots referred to by content and these snapshots that can reference other snapshots to form a DAG. Just because you didn't take one hour to understand this doesn't mean it's "complicated".
> That said, I LOVE GitHub and gitlab.
These would not have existed if not for Git's internal representation that allows the necessary flexibility for other tools and extensions to integrate directly.
Worst is when they're data-compressed so you're not only pulling all historical versions but also most every byte of each historical version.
I wrote an article for the Fossil docs covering this, but the core experiment can be run against any VCS:
https://fossil-scm.org/home/doc/trunk/www/image-format-vs-re...