SQLite Doesn't Use Git
matt-rickard.com
matt-rickard.com
Keep in mind that Fossil was written specifically to support the SQLite team's workflow. It was never designed for _anyone_ else's needs. They don't care if it doesn't work for other people's work styles because they didn't write it for other people.
Yep. One thing I like to ask is, does your project have millions of lines of code, being worked on by thousands of developers scattered around the globe? If not, you might want to think about whether Git is really the best fit for your needs.
Apart from its ubiquitousness, these are all neutral features: there are lots of crappy piano players out there who might be better off learning a different instrument, but if you don’t have any other preference it’s hard to go wrong with the default choice. Git is one of those tools where it’s really not worth the trouble looking elsewhere unless you know what you’re looking for, because there’s value sometimes in going with the flow. Whether it was for the best that git was the tool we all settled on, I think there’s some good in the fact that we settled on something.
The right question is why did git need github in the first place (while svn didn't)
It legit may be a few % better for a few workflows, but why die on that hill?
Is that the same Mercurial that didn't even supported stashing local changes?
Exactly what does anyone stand to gain by switching from Git ti Mercurial?
But it’s a good example of why mercurial was better. It’s more ergonomic. All of the parameters on shelve match what you’d see on other hg commands and to shelve and unshelve you use those commands, not pop and push. You can also give short names to shelves for use in the ui, not just the message or sha.
Not really. Mercuria's shelve extension existed for a while but it only saw its way into Mercurial Core and made available by default with the release of Mercurial 5.1.
https://www.mercurial-scm.org/wiki/ShelveExtension
I feel it's disingenuous to clam that a feature was provided by an application if it's only provided as an add-on that requires the user to explicitly enable.
The standard way mercurial added new features was to gate them behind configs and only after they were widely tested turning them on by default. This is a good thing!
Prior to shelve being included with the mercurial distribution a third party extension was available going back to 1.x (or perhaps before I started using it around 1.8). And get this, it also prioritized ergonomics and was largely just adapted directly.
Would you consider it disingenuous to claim that there's visual UIs for git? It seems weird to ignore extensions that are extremely well known within the community. Plenty of people recommend using Firefox because it supports better ad-blocking extensions, even though those are also "sold separately".
Git stash however, is one place where it's just crap. `git stash apply stash@{3}` ?!?!? No. omg no. Sure, for "hey it works" but not for "hey let's have this be the standard interface that we give to everyone and never improve" After all these years we still can't say `git stash apply 3` ??? really?
also git stash is wildly divergent in terms of conventions relative to all the other commands.
Sometimes decent software accidentally becomes a ridiculously broad standard. You can't resist that force by dismissing it as what it was originally meant for.
It doesn't matter if git is the best fit. ~Nobody under the age of thirty is picking between git anything else any more than they're picking between Linux and anything else to run their single node/java/postgres process on their prod servers.
This personal assertion makes no sense. What leads you to assume that just because an application scales beautifully that automatically means it doesn't fit your needs?
I mean, you don't even bother trying to argue feature X is disappointing or feature Y is missing. Your whole point is that git scales so if you don't need scale then... Then what?
The truth of the matter is that Git works fantastically well both with personal static website projects with a dozen files and with million loc behemoths, whether in local repo only or with multiple remote repositories.
I also agree "distributed VCS" like git may be more complexity than many teams need. I can't remember the last time I worked on a repo offline (arguably the signature feature of distributed VCS), ever, although I can understand how this can be hugely helpful for some teams.
IMHO, this solves all issues in a very elegant way - my local repository shows all the commits (both private and public) but everyone else's shows all the commits as intended by the developers. Additionally, the public history is maintained consistently without wondering if someone messed it up.
https://www.fossil-scm.org/home/doc/4d8aecdf/www/private.wik...
Also, do I understand it correctly that it's a single branch named "private" per repo?
Git style mailing of patches?
"Who decides where one level of granularity stops and the next begins? I think it’s the author of the commits. My workflow over the last ten years is based heavily on being able to massage commits so that I can prepare what I share to the server repository, where it can no longer be changed. I agree that there should be an unalterable history, but disagree with the author on where that history begins."
"My workflow is considerably different than it was before I used Git or had access to rebase. I would now be much less efficient if I didn’t have rebase. It would make me constantly focus on cleaning up commits before I really care to. You could make the argument that cleaning up afterward takes more time, but I haven’t experienced that to be the case. Instead, I want to be able to set the priorities rather than worry about committing something that I cannot undo."
You don't need complicated branch refactoring tools when you're only sharing with a handful of people. It's like working with RCS back in the day: history is still linear and no one needs to worry about cross-ports between feature branches.
[1] Ironically done using a github mirror, because those are the tools I know and they're fast.
5 contributors, with over 70% from a single person is indeed small.
It so happen that I usually do commit squashed changes to many files over code base with long and heavy legacy. I am afraid that squashed changes lose their utility for bisection search for a bug.
Yet, I required to do so. Mainly, because I can.
I don't have a problem with squashing per se, especially when it's a private branch for a single developer. The problem is when you have multiple developers working on a feature branch together, or when you have to merge in a branch with a fix or feature that has not yet been merged up higher.
My bottom line is I want as few merge conflicts as possible, and I want to see the real author and a relevant commit message if I'm digging into changes to figure out what went wrong with something.
It tends to be developers that are used to the opensource workflow of forking before submitting a patch or pull request. Who are now on a team where everyone's working on branches in the same repo. I understand wanting to combine a bunch of commits that say "trying it like this, trying it like that". but I find it's better to have the noise than lose something important, or cause a messy conflict later on.
I also don't like multiple changes in a single commit. I hate when I look into the commit history of a file and it's "added feature X" which touches 15 different files making changes to multiple code blocks in each. If I want that, I'll look at the merge commits.
Usually, private branches are get deleted after merge happen.
In the end, I will have a situation where there is a defect in the main branch and there is no way to bisect it narrower than to squashed commit.
What should I do with 3200 changed lines in 320 files when there is a defect?
[1] https://stackoverflow.com/questions/4049958/embedded-softwar...
The page above hints that pre-review defect density for C/C++ code is 4.5 defects per 100 lines of code. After review it is 0.82 defects per 100 LOC, or 8.2 defects per 1000 lines of code. My big patch will introduce 25 defects that'll pass review. Even 10 fold decrease in defect density after testing (haha) will not significantly reduce probability of the introduction of a single defect.
Either you can break your big commit down into pieces that can be applied one by one - then do that using history rewriting (and testing each part!).
Or that's not possible, e.g. because it's a global architecture change or something that is already atomar. Then there is nothing to bisect anyway - there is nothing that can be done, no matter which tool you use and no matter if you squash or not.
Having private branches also does not always work. For example, there was a purge of unnecessary branches in our main repository. If I use a fork of main repository (and I do), my private branches are also not readily available to the team.
The choice is done for you. It is not yours or theirs often enough.
Personally, I think that the only bar that is necessary is that commit history should remain readable - which is subjective, but that's why we have code reviews. In any case, this is not an issue with the tool, but rather with the religion that formed around it, which is often the case with software tooling.
One thing that helps a lot with a pragmatic consensus is to ensure that all members of the team have personal experience debugging old code that they didn't write or previously review.
Fossil DOES allow you to modify history, by adding new artifacts which retroactively change how previous artifacts are interpreted. They are called Control Artifacts and the set of things they can change are limited, but there's no reason additional ones cannot be added.
> Fossil does not support rebase. Here's an article from the author titled, Rebase Considered Harmful. While Fossil has the ability to squash merges, the primary workflow supported is merging.
In particular, they show the following comparison:
Merge:
-> C3 -> C5 -> C7
/ /
C1 -> C2 -> C4 -> C6
Rebase: -> C3 -> C5] -> C3' -> C5'
/ /
C1 -> C2 -> C4 -> C6
However, a more realistic example of a feature branch that lives for even a few days looks like this:Merge:
-> C3 -> C5 -> C7 -> C9 -> C11 -> C13 --------> C15
/ / / / \
C1 -> C2 -> C4 -> C6 -> C8 -> C10 -> C12 -> C14 -> C16 -> C17 -> C18
The final history after rebase being: -> C3' -> C5' -> C9' -> C13'
/ \
C1 -> C2 -> C4 -> C6 -> C8 -> C10 -> C12 -> C14 -> C16 -> C17 --------------------------> C18
A much much simpler history to follow, and obviously shorter.And the complexity without rebase grows even more if more people are doing local merges and pushing to the same remote branch.
In fact, following the same merge history as in the first example, I should have named the final ones C3''', C5''', and C9', since they get re-written several times. But, most of the time, this is entirely irrelevant to their history (especially when they don't even touch the same files as the changes being merged in).
Note that C7, C11 and C15 completely disappeared, since they are unnecessary. Instead of being extra commits clogging up the log, they are the history rewrite events that don't need to be consigned.
In my experience, there's no such thing as an "unnecessary" merge. Every merge commit encodes tree changes from the various merge strategies. Explicit merge commits keep those (mostly) separate from deliberate developer changes and make it easy to trace back a subtle merge bug or the wrong merge strategy.
Git's modern merge strategies today feel almost like magic and "never mismerge", so I don't blame anyone for never having had to trace through merge commit history for subtle mismerges today, and assuming in all cases the merge strategies "just do the right thing". But it still happens, those subtle, terrifying moments when the magic breaks down in subtle ways and the merge tree output isn't what you expect or need.
From my perspective, trusting rebases and squashes and cherry picks is putting sometimes way too much trust in git's various merge strategies. I do trust them, and rely on them for a lot. But I've also seen the dark sides of them: the need to find mismerge needles in large haystack branches, the gut wrenching low level micro-management of rerere caches to avoid similar mistakes in the future. Personally, I'd much rather have a million "unnecessary" merge commits than ever again need to mentally "bisect" a mismerge from a deliberate choice in some long gone developer's hand cherry picked and squashed commit where all the merge details are mixed in with all the other code changes and no clues where the dividing lines were.
I can imagine that sometimes a feature depends on changes from develop after the start of the feature branch. This means that development started too early.
Cherry-picking my changes on to a new branch will hide that fact. Rebasing will conserve the authoring timestamp. I guess it depends on the specific situation what's best.
In a trunk-based development workflow, I'd have to merge asap and maybe hide things behind feature flags or compile-time switches to avoid such merges.
Things get even more interesting when the merges from develop affected files that were changed on the feature branch. You will likely get conflicts if you attempt to rebase them away.
Rebasing your work onto the head of master before merging it in, versus cherry-picking the relevant changes from your branch onto the head of master, are perfectly identical in terms of end result - whichever workflow is easier for you will achieve the same thing.
And of course you can get conflicts if multiple people are developing over the same files - this will happen regardless of using merge, rebase, or cherry pick (or even patches over email). You fix the conflicts when merging / rebasing / cherry picking / accepting a patch.
They do if yet another tag or branch still links to those commits. And in practice they'll stick around for a while anyway, accessible through the reflog.
Fossil requires users to either commit stuff they don't want to see later, or hack away until finished, cherry-picking changes as necessary. This isn't a good workflow for everybody.
Git allows users to do anything they want, but your Linux master branch may not be the same as Linux Torvalds' master branch, and history can be changed at anytime, while merge commits can be completely rewritten, losing the actual merge information.
Both of these are wrong.
There should be a way to make WIP changes without committing it nor committing to it. (Yes, those are different things.) But there comes a point at which a change should be committed and committed to, to never change again. (With the exception of the nuclear option of erasing accidentally committed secrets and the like.)
Fossil is too rigid. Git is too flexible.
There is a middle ground. My design for such a VCS would have two kinds of commits: saves and full commits. Saves would be WIP stuff; anything goes. They would also be on auto-generated branches. Once WIP stuff is done, it should be to be rebased, squashed, merged, and/or cherry-picked into actual full commits on the branch that the saves created a branch off of.
But once those commits have been created, they are there forever. They should be a commitment. (I love the double meaning of the word "commit" here.)
With this design, people could have their cake and eat it too: they could have WIP stuff that they can still massage into a history they like, but then they would also have solid, reliable history with proper merge commits.
I kid, I kid! But I do commit just to push a back up before getting on a train, or a convenient point to diff against as part of a refactor, etc.
They should all be a single commit by the end.
If squashing happens then all I can immediately say is "something in this 600 line changeset over 12 files for Feature X broke it". The bug report for that one is going to be more vague, get allocated more story points, and maybe stay in the backlog for several sprints (or forever).
If people are pushing their feature branches and not deleting them that makes life a bit better, but generally people who want their git history to be "clean" and squash commits also want to get rid of old branches.
Someone might say code review should have caught it. Code review almost never catches actual bugs. Or tests? Tests only test stuff that the author expected to break.
The main things I try to ensure are that every commit is small, every commit has a useful message, and no commits should break the build.
Personally, I prefer my coding to be free of such distractions, and to use SCM facilities as a scratchpad for fast iteration on the code (the ability to easily reverse changes etc); this results in many commits that don't make sense in the final pull request / code review.
The bulk of my PRs are still one small commit, but when a feature gets larger I find that "what's the next thing to try to solve the problem" isn't always the same question as "what will make this most readable for a reviewer." I tend to address presentation of the commits in the PR at about the same time I'm addressing readability of the resulting code, as they wind up being related concerns (though not exactly the same thing).
I can't think of any special tools you'd need for this? A git history visualizer with the ability to easily diff and squash ranges helps, but this is nothing special for pretty much any modern IDE. At the end of the day, you just look through the WIP branch history and identify chunks that span multiple commits but logically represent a single change towards some goal.
In my view, it is not subjective. All commits locally must pass unit, integration and e2e tests, and be locally tested by the developer, every...single...one.
I'll have to try it sometime, I have a feeling it wouldn't normally turn out too differently to what I end up with anyway. I do think all commits should build and run properly - otherwise git bisect and similar processes stop being useful.
WIP of a function/method, they could be squashed into a good commit that explains what and why. But some people implement a full feature, squash all of it into a single commit that simply says 'implemented Y'. But you loose context on why they modified the different functions/methods.
Yes, you can try to find it, but there's no context so you're just dealing with a huge block of code that modifies other things, doesn't just add new code.
You don't need to have things tidy, you know you can recover anything you do. So likewise when you're "done" you can squash and split them into changes that make more sense logically before pushing to some shared tree that someone else is going to have to reason about.
People who develop without tools like this tend to do it in a big flat directory and think about "commits" as something done every day or so. Once you get beyond that style, it feels really clumsy.
Which keeps context which is my point.
The other case that I see is squashing it all into one that simply says 'Implemented feature Y' which doesn't provide any context into why something was changed.
If you don't backup your code to remote even if it's not finished, you're taking serious risk.
Of course it's a branch other than master/main, even if it's on "remote".
The thing is, I do not always continue developing on a single/same system, so I do "transfer" commits. I push unfinished code to its own branch on remote, pull from other system and continue developing.
When the code completes and passes all the tests. I merge to either development to prepare for the next version, or to master if the utility is small enough.
The place where the development branch lives doesn't matter, and --amend ing a single commit during a feature development is not the most correct way either.
Oh, and I don't use GitHub for my own software. That part is over.
The drawback of one remote endpoint is that it becomes less obvious whose rules apply. You could argue that other people can mess up your remotes, but it isn't until you get to very large organisations or public development (e.g. FOSS) that you should need to deal with adversarial behavior.
I've had bosses who got angry that I force-pushed to my own remote branches because they liked to review code without being requested by pulling the branch, and after a force-push that doesn't work. But my defense was always: Let's establish a protocol on how to cooperate; there is surely some way we both get what we want, and that git supports.
I mostly avoid creating these kinds of conflicts when people's limited git experience would cause stress or unnecessary use of people's time fiddling with the history. Or if more than one person is doing stuff independent of one another.
I've had colleagues that frown on rebase, and people who can't not.
One of the beauties (and complexitties) of git is that it allows for this diversity of workflow.
At other time I make "transfer commits" since I'll continue development on another machine and need the latest snapshot of the code to continue.
Not all features fit into a single new function, we need to move mountains to rearrange stuff, and to keep code tidy.
As long as the commit messages are clear, and the code works at the end, it's alright.
Of course. But when modifying another, you can commit and say,
Implemented function X. Modified Y to accommodate this new argument that's used on X... and so on
>As long as the commit messages are clear, and the code works at the end, it's alright.
I'm fine if the commit messages are clear. The problem is when squashing, some people squash and don't keep commit them clear or some even squash into one that doesn't provide any insight or clarity, 'Implemented feature Y' and you get a diff of thousands of lines that touches everything
Sometimes I need to write "Implemented function X, but evaluates the result wrong possibly because of this. Fix this first, then continue".
> The problem is when squashing, some people squash...
A single commit touching whole codebase and only says "Bug fix" (or similar) is bad. I concur.
Also commit history should have enough granularity allowing bisection and partial rewind to understand problems and other side effects.
Code for function <B> has the skeleton. Still need to add <...>
If that's how a team develops PRs, then the commits are structurally suboptimal, and the problem is in the development practices.
Such team won't be able to efficiently bisect, independently of the repository structure.
I personally work with granular, self-standing commits, and with this workflow, non-squashed commits makes sense.
Obviously it's not possible to always have self-standing commits, but it is possible the vast majority of the times.
It takes a lot of practice and discipline, though, and if a team is not willing to put them, then of course, squashing in the only way that makes sense.
Edit: I strive to keep my feature branches both short (# of commits) and light (# of lines changed).
If the branch was squashed, bisecting will be less effective, compared to the same branch, merged without squashing (at the conditions of the commits being self-standing).
If this title were "SQLite uses Fossil", it'd be more apparently just another ad for Fossil.
He likes to "engage" his audience in a way.
From my observations, some people think that they can extract more arguments or have a more fruitful discussion by being slightly spiky like that. I do not share this view.
If i remember correctly, it wasn't that they had any real problem with git, but that git didn't really fit with how they wanted to work, so they created their own thing.
SQLite is what I consider a well-engineered product.
And Richard Hipp is an excellent engineer. Listen to his podcast episodes on the changelog. It's fun. His mentality is interesting.
Fossil has the ability to push to git, which I do periodically just for curiosity.
An example would be the excellent diffing tools accessible through localhost, as well as things like issue tracking, which git doesn't understand as a concept.
There were some deal-breakers for me (kind of esoteric) and I switched back to git, but not without some annoyance.
I run an OrangePi Zero with 512MB of RAM as a home server since it both works and is a fun experience.
The gc command doesn't help, and git automatically runs it after some operations.
I use Windows + Linux machines + Pi and share between them so don't use the Pi4 exclusively but have noticed a difference in its ability to handle and open the repo.
Ever try and copy the .git part of a git repo? When it gets big, the copy gets bogged down by the sheer number of files.
All of this is also possible to do on a repository locally, without cloning.
This comes about from their development styles (see: http://www.catb.org/~esr/writings/cathedral-bazaar/).
- Use fossil if you prefer the cathedral style of development SQLite is open source but does not accept contributions.
- Use Git if you prefer the bazaar style of development. You want to accept hit-and-run style contributions from anyone.
If there’s commits like “fixed off by one issue in refactor” then it should be a comment in the code so that future will explorers won’t make the same mistake.
Rebase by itself isn’t bad, it just has to contain a good description of what was done and why.
Edit: after posting I realize that I confounded rebase and squash — which is kind of ingrained in my head as that’s how I use GitHub to merge PRs.
It's like a kind of blame + log in a single UI, in Git terms, and is just wonderful for quickly moving through the entire history of a file.
There is no consistency in the language or syntax of the commands when taken as a whole. Even individually some of the commands don't really make any sense and harken back to seemingly random unix CLI flags. "I wasted enough time to memorize them" doesn't mean they couldn't be cleaned up and made a little better.
If you're a single developer git is indeed overkill (not that there's a much better alternative).
If nothing else, network effects. Linux uses git, github is popular.
Unpopular opinion, but I find the value of having the actual surrounding context in which the code was developed more valuable than having a flat history.
A non-linear commit history is much more difficult to run a bisect on which was enough to sell linear for me.
That's not to say each PR is one commit, if there are multiple distinct parts of a PR that could be cherry picked in a way that they make sense on their own they'd still be separate commits.
This thread is about squash merge, which converts the whole branch into a single commit.
But only if those commits are real atomic changes of code which compiles and not a series of 'oops', 'tests pass once again!' , 'forgot to commit this file' type inner-monologue type history.
And then there are shops where less technical authors are contributing solely via the GitHub WebUI. So you get a commits of the ilk 'Updated <Filename>' or ''Files added via upload'.
there are lots of people who use emacs for those two things and other editors/ides for other things.
i use emacs for more than the two, but certainly not for everything.
I'd love to hear some real-world experiences.
Like SQLite, Fossil is very nice software.
In what possible world is telling the most successful set of design decisions about document management in the history of documents or management to get fucked not just “I’ve got goodwill to burn, let’s ride.”? SQLite is indeed badass, but cool it djb, I’m already stressing about the zero bugs ever thing.
To squash or not to squash is a never ending debate precisely because of that. There isn't a right way to do it, each approach has a set of tradeoffs.
Maybe they are reinventing the wheel, but at their scale it's a tiny decision that doesn't mean that much. If they find out they are wrong, they can choose to migrate to Git instantly.
SQLite doesn't solicit contributions from programmers outside their organization, so it's really no skin off your back.
Edit: I’d like to apologize for mouthing off like that. I’ve earned some amount of cranky old-timer cred, but not enough to just be a dick, which that was.
Acme is great. I haven't been able to find another editor that integrates as tightly with the operating system. The closest I've gotten is Kakoune. And I wish mouse chording were a thing more generally. If we're going to use mice, we should optimize their utility.
I easily hit 4hz at the end of a 20 hour coding bender, and I’m old: the hot shit kids on modafinil must be doing 1.5x my decrepit ass.
If you can really move on the keyboard? Touch the mouse? I’ll grab a coffee while I’m at it.
Edit: It turns out that “fast emacs” doesn’t turn up a video of a real pro easily on Google. You can see a code God like Russ Cox being slow and clumsy as fuck if you want: https://youtu.be/dP1xVpMPn8M.
It’s a good thing thing he gets everything right on the first try (not sarcasm, Cox is an alien life form optimized for doing the impossible) because if he didn’t he’d finish pound-defining everything by about next March.
Great read :D