An overview of version control in programming
lemire.me
lemire.me
In some circumstances it's helpful to think of commits as sets of changes, but Git's commit are snapshots and not changesets. Converting commits into sets of changes is done on-demand.
But then you have cherry-pick, which seems to work by "applying a commit" (you can't "apply" snapshots...I think) with git cherry-pick hash. That's the same thing as git cherry-pick hash^..hash which seems to more clearly indicate that we're applying the diff between two commits. And rebase, which really does present itself as a tool for reordering diffs.
Anyway, everyone knows what you mean when you say "apply this commit" so maybe this is just a stupid complaint about "consistency" with no real point. Especially since git does use both deltas and snapshots under the hood.
One thing about Subversion that I miss in Git: when you move a file or a directory, Subversion tracks where it came from. That's not true in Git, which just stores snapshots (mostly) and has to guess whether a file was moved. That means that `svn log FILE` behaves like `git log --follow FILE` by default (except always right rather than "right when the heuristics work"), which is much more useful. One thing I dislike about Git is that its history model tends to crumble when you move files and directories around.
All in all, it’s a hard problem and there’s no perfect answer I think.
1. Move the file verbatim.
2. Change the file as necessary.
Subversion will remember the move because it always does — Subversion stores changesets internally, and those changesets contain history metadata. Git won't track the move, but when you run `git log --follow` later it detects that the file was moved because the content of the file that was added was the same as something which existed before.
It's a little awkward if only a small portion of a file was migrated, though. In such a case, you may just have to fall back to documenting in the commit message that code was moved.
This approach has another benefit: it is easy to review. The first commit can be validated as "verbatim move — check!". Then the logical change can be validated with effort proportional to the amount that was changed. In contrast, if you lump them all together, it can be hard to discern what changed when presented with a huge diff containing mostly verbatim move but also a few subtle changes.
Git's form of compression is very interesting, but a key concept as you study the history of version control is that snapshots and deltas encode the same information, so the choice of which to use is an implementation detail[1].
[1] I covered this in detail in a preso on Git data structure design for Papers We Love San Diego: https://www.youtube.com/watch?v=fHSZz_Mx-Uo&t=400s
Svn has some warts (shelving feels sooo half baked despite having like 3 different ways for that) but at least reverting to an older state isn't one of them.
Not that I expect anybody to maintain historical perspective. Linus Torvalds famous 2007 rant at Google shitting all over Subversion and its developers set the tone for a long tradition of anti-Subversion anti-history. Never mind "learning from what's gone before", it's like there was only backwards progress until the tool du jour was suddenly invented ex nihilo.
Revert to the previous version. Push as new version. Fixed? SVN was used widely and for years, prior to git. I remember the biggest problem being expensive (full copy) branching.
Copy on write is an expensive operation for branching. SVN repos that are multiple gigs, in size, are a pain to create branches for.
What is slow in SVN is switching to the branch client-side, because branches are always created server-side and you still need to synch the whole contents of the branch from server to client when switching (which works similar to a regular `svn update`).
"svn cp http://.../trunk http://.../branches/foo" may not be the most concise command to type, but all it does is create a new item "foo" in a new revision, which references a previous item "trunk" in a previous revision. How many files/directories are under that doesn't affect how long the operation takes.
https://www.joelonsoftware.com/2000/08/09/the-joel-test-12-s...
That said, if you're worried about how your questions are perceived, a little rewording can suss this out and give you more context. I like to ask:
What does your software development process look like? Tell me about how you manage, test, and deploy your product.
You can rephrase a lot of questions in a similar manner to get what you want and leave room for the interviewer to expand (or justify!) their responses.
Usually it's sigh and a nod of the head (as if they are remembering THAT company). That's a good sign for me.
There is a mutual recognition that we've both worked in some terrible companies and learnt how NOT to do some things. This generally means (and some good followup questions) that they have CI/CD,a PR review process, some ticket tracking, agile/kanban etc.
I'm generally not worried about how my questions are perceived as that's a red flag for me. If they have a problem with that, then most likely not a good place (culturally) I want to work.
Perforce has something similar too: http://ftp.perforce.com/perforce/r16.2/doc/manuals/cmdref/p4...
svn merge -c-123 . # single commit, the second "-" is "do the merge backwards"
svn merge -r123:122 . # Multiple commits in reverse order, "the change going from 123 to 122"
https://stackoverflow.com/questions/13330011/how-do-i-revert...That was the promotion.
In fact RCS was never as fast as SCCS. RCS did not, in fact, use less disk space than SCCS. It is hard to identify any particular where RCS was better. The code quality was abysmal. But Tichy got a PhD out of it, so there was that, anyway.
I’m curious about the differences between git and mercurial, are there benefits in choosing one over the other?
It is coded in Python, which should have made it slow, but it was very fast. All the data motion and analysis bypassed the Python interpreter.
There was another called Monotone, where Git lifted its data model from, wholesale.
And there is Fossil, which is growing in popularity. If it offered a way to squeeze intermediate edits out of the revision history, it might grow faster, but the author is hostile to the concept. Where it really shines is in never corrupting its data store. I have had Git corrupt its data store quite a few times. By luck, I have not seen it happen on a server others relied on, just on my own cloned repositories. But the tricky stuff is mostly done to cloned repositories.
Are you saying when a private branch is published, it shows up in the public repository as a single diff?
Mercurial has immutable history, so no squashing commits, no deleting branches, at the time I was using it there was no amending commits, in fact reverting a commit doesn’t even come enabled out of the box! Some folks loved it; no changing history, everything documented as it happened. We practiced trunk based development, so no branches except for hotfixes, so there wasn’t a lot of sprawl.
Ecosystems largely don’t support mercurial, so that’s definitely a consideration. Since Merge Requests and feature branches are largely practiced now, I feel like there’d be a lot of noise in a repo if folks used mercurial.
I don’t particularly miss mercurial, personally. I’m less into “pure” workflows and forcing behaviors. I think git is super flexible and generally practical and I’m overall pretty happy with it.
The history editing capability of Mercurial is arguably more advanced than git, especially in a collaborative setting, because of Changeset Evolution [1], and the Evolve extention [2]. The former keeps track of metahistory of commits. The former keeps track of the metahistory of commits, and synchronises it between repositories. The latter provides a set of expressive command line tool to edit history. With them, collaborative history editing and stacked PR is a pleasant experience.
[1]: https://www.mercurial-scm.org/wiki/ChangesetEvolution [2]: https://www.mercurial-scm.org/doc/evolution/
- evolve [0], which allows to rewrite history lossessly (without ever risking losing data)
- absorb [1] which takes uncommitted working copy changes, and for each hunk finds the last commit that touched those lines, and rewrites it. It's an extension originally from Facebook, in core since 2018. Works like magic: no "fix" commits ever more.
Plus, all of this is available using mercurial locally and interacting with git (and github) remotely, via hg-git. Admittedly, this requires to be a bit of an advanced user, but the gains in ergonomics are tangible.
[0] https://www.mercurial-scm.org/doc/evolution/
[1] https://gregoryszorc.com/blog/2018/11/05/absorbing-commit-ch...
>The public phase holds changesets that have been exchanged publicly. Changesets in the public phase are expected to remain in your repository history and are said to be _immutable_
Note that in Git all history is mutable, even if published.
Regarding 'backout' (and 'revert'), to the best of my knowledge, it does not revert commit, it creates a new one (reverting changes), and I frankly do not know is that's possible at all to amend commit in Mercurial (when I worked with it, that was definitely not possilbe, but that was a long time ago)
hg backout is like git revert. Creating a new commit is correct if you want to e.g. propagate the change through continuous deployment.
And hg commit --amend has existed for a long time.
For amend option - well, we switched to Git at time of Mercurial 2.1 (I said it was long time ago), did not notice they added this feature, sorry.
> hg commit --amend
to change the topmost commit.
Marking commits as public is mostly a safeguard against accidentally altering history that others may already depend upon. This is just there to provide awareness of the giant footguns hiding when editing history after it has been shared (git contains the same footguns without safeguards). You can revert the status of a commit from public to draft and then change it. Just like in git, it's very dangerous to do so, but hg makes it very obvious. The command is
> hg phase --draft --force .
On the other hand, if you're on Windows, TortoiseSvn is one of the best VCS interfaces I've ever used.
- It isn't yet another tool which shows you all your directories and files (in addition to Explorer and your IDE/editor).
- As Subversion was always designed to be a library (as it was already obvious then that VCSs would be used from other tools like IDEs) the integration is much better than TortoiseCvs or TortoiseGit which compose command lines and then execute them.