The Biggest and Weirdest Commits in Linux Kernel Git History (2017)
destroyallsoftware.com
destroyallsoftware.com
When it came time to pull things together, we thought it'd be fun to try and maintain the histories. So I added his repo as a remote and simply merged his unrelated history into mine.
Fast forward and now we offer hoodies as swag to anyone who contributes to the repo. We personalize the hoodie with your git username, the truncated commit hash of your first commit, and the number of parent commits to your first commit.
Having 2 root commits means that both my cofounder and I have hoodies with a large 0 as number of parent commits. Just a nice way to commemorate this accident of history :)
It makes for a great conversation piece, even around non-tech crowds.
Has been particularly interesting to balance the different languages in the same repo. Few painful bits, but overall has worked out really well.
We also build the entire thing into a single Docker container containing every service. You just start the particular one you want via the command (the entrypoint is a little shell script to invoke the correct executable).
Otherwise if you ever need to do a rollback or post mortem or similar you have to figure out what commit out of the 20 different projects was actually being used at the time.
Of course that comes at a cost and there are other ways to get the same benefit, but my experience is that this is an easy way to keep things consistent.
- much less admin, don’t need to maintain/update so many repos
- easier to coordinate changes between systems at the same time
- everything is on a single version of each library - you don’t have some services on one version, some on another, etc
Benefits of having a single Docker image:
- much faster (overall) build times - doing everything in one go is way faster than doing each one individually
- much faster deployment - only have to download the image once per host, rather than many variations
- most services have a similar set of the same base dependencies, so reduces duplication
- overall the difference in image size between including one service and including all of them is probably only 5-10MB anyway
Those are just the things which immediately come to mind. There are probably many others.
Image layers are immutable and get deduplicated/cached individually when stored or sent over the network. For example if you pull 10 images with a common base to the same host, the first image might be slow to download but after that docker only fetches the deltas between it and the rest.
In fact, for at least one of our components, we've already done this. We had something that solved the same problem for a couple of different permutations of software versions and platforms. The person who had originally written it had used separate repos, with common code copied and pasted between them, and later on we merged the repos together and factored out the common code so that could be shared. So we already have at least one repo which has 4 root commits, I believe.
Actually, I just checked and it has 7 root commits. I'd forgotten that we had also had similar problems that needed to be solved for some other software, and had some interns who had done those in separate repos as well, which also eventually got merged in to share all of the common code:
$ git log --max-parents=0 --pretty="format:%h %cd %s" --date=short
666c0bf 2012-07-09 Initial tree.
528ca3a 2012-06-12 first commit
50f1b18 2012-06-12 first commit
86d66e0 2011-11-17 initial
76a7789 2011-03-21 initial commit
58f7afc 2010-05-20 initial
deb50b7 2009-10-29 initial commithttps://imgur.com/gallery/oiWeZmm
Run this:
git log --graph --abbrev-commit --decorate --date=relative --format=format:'%C(bold blue)%h%C(reset) - %C(bold green)(%ar)%C(reset) %C(white)%s%C(reset) %C(dim white)- %an%C(reset)%C(bold yellow)%d%C(reset)' 2cde51fbd0f3
"everything is linear if plotted log-log with a fat magic marker"
While technically git is great, usability wise its not (for me).
I've joke it's the source control version of stockholm syndrome before.
Anyway, my comment is that I still need to use CLI for some things, ie. Git merge squash, so while the GUI apps are good learning tools for git in general, you still need to Google things for CLI consumption if the GUI doesn't support it and you're not a git expert.
This jab at Linus is unfounded, he have replied calmly and professionally.
Wow, that's a really charitable way of spinning them.
----
Those who only know Linus from his rants might be surprised that here "his usual fashion" means:
- Acknowledging that the root cause was Github's documentation being misleading.
- Not blaming the contributor for being mislead by Github: "I can see why that documentation would make you think it's the right thing to do."
- Admit that the ease with which the accident happened is a deficiency in Git's UI.
- CC the Git maintainer to discuss improving Git to make it harder to do this by accident. (Which eventually lead to the --allow-unrelated-histories flag being needed to do this kind of merge.)
Granted, there are some jabs at the user who made the mistake:
> Why did Laxman do that insane merge? Why did it get
> pulled back?
>
> You actually have to work at making shit like this, so I wonder what
> workflow you guys had to make that bad merge.
If you wanna put more spin on this apparently this merge came from Nvidia.
Manishearth, a Mozillian and Rust contributor, wrote this post on splitting apart a repo. He's good at explaining all sorts of things.
https://manishearth.github.io/blog/2017/03/05/understanding-...
This made me chuckle.
I’ve built a fairly simple tool that tracks which branches to merge into my beta branch and automates its regeneration, and I’m polishing it up so it is distributable and can just be a regular Git subcommand, git-managed-branch. What I have already is useful to me, and I think it’d be useful for many others as well.
The essence of the work of the script I have at present boils down to this:
git stash # (if necessary)
git checkout --no-track -B origin/master staging
git merge --no-edit feature1 feature2 fix3 fix4
git push --force
git checkout -
git stash pop # (if necessary)Does anyone know? (Am on phone, slightly difficult to check).
--allow-unrelated-histories
By default, git merge command refuses to merge histories
that do not share a common ancestor. This option can be
used to override this safety when merging histories of
two projects that started their lives independently. As
that is a very rare occasion, no configuration variable
to enable this by default exists and will not be added.This reminds me of the last interaction I had on #lesswrong - I openly speculated about the UI of a system I didn't have access to, because I wanted to chat about it with people I generally respected. But some of those people just weren't in a good mood and lambasted me for not simply "googling it for the answer." Never mind that the answer isn't easily found, and never mind that that completely invalidates my whole reason for interpersonal interaction anyway...
Auditors LOVE immutability. To be fair, git doesn't provide that, but it provides the next-best alternative: tamper detection. If anyone rewrites history, git will show that. The gitrefs between two points in time will not match if anyone has modified data or commits in the meantime. The auditors also have no problem looking at previous years' documents where they have recorded the relevant gitrefs at the time.
This has gone so far that this year's policy review was a breeze. Our compliance documentation is maintained in a git repo, with all documents as markdown files. The final documents are simply compiled PDF and HTML artifacts.
In 2016, the auditors asked if we can provide snapshots of previous policy versions. In 2017, they already understood that we have everything in git, and knew to ask for clarifications as to when a particular change was done and who had signed it off. This year our auditors literally asked for the latest compliance documentation bundle from CI, all the individual commits, and the overall diff over the year.
Wall time spent for policy review: ~20 minutes.
(Understanding that some effort will need to be put in)
The important thing is to never treat compliance audits as box-ticking exercises. That's a never-ending, vicious cycle. In fact, many of the findings are simply different aspects of the same thing. You can pre-emptively work on this: identify what parts of requirements are essentially duplicates, and make improvements that satisfy all of them at once.
Then proudly flaunt them. When you can show to the auditors, in person, that you have considered the wider business implications and worked to understand the compliance requirements, you are on much better ground. That buys trust.
Then, educate the auditors when necessary. Show them in practice how something simple can provide a better trail and an improved experience. When possible, provide evidence in the format they initially ask for, but also in the format which is more suitable and more convenient. Auditors are humans. They just often are not aware of the leading edge, of what is possible. Show off solutions that are more convenient to both of you.
They will learn. They will be impressed by some things you do. Anything that makes their job easier, while satisfying the intent and spirit of the audit, will be an easy sell. They also believe in repeat business. Show repeatedly that you know what you are doing, and why you believe that your approach makes more sense (while delivering better audit trails).
Convenience is a strong currency.
0: There are some domains where a single company has essentially a "Royal Charter". These are much, much harder to deal with, because the monopolist has little need to employ personnel with proper technological understanding. They can also strong-arm and bully their customers at will, because there are no alternatives. Audits like these can very easily degrade into box-ticking bonanzas. My advice for these cases is: pick your fights. Double down on what you truly believe and give in on smaller, less disruptive items. Rinse and repeat on subsequent years.
I find this looks a lot cleaner in tools like Gitlab/Bitbucket as well.
I don't think you need to explicitly check out a new branch and then cherry pick.
The best approach (which is what git.git itself uses) is to do the work on a separate branch, which is rebased often; once it's moved to the master branch, it's "frozen" and won't be rebased anymore. All pull requests are based on the master branch, so aren't affected by the constant rebases on the development branch. (Actually, all the work is done on topic branches, what's often rebased/rewritten is a sequence of merges of these topic branches into the master branch.)
You can give people access to the staging repo, but then they have to acknowledge that they are willing to deal with some occasional chaos.
It worked well until someone ran an accidental `git merge` of one of the submodules into the main project which sewed a bunch of confusion. After that point the practice was banned where I worked. Nice to know that git added a flag to prevent that, though at this point git have also added a subtree command which I think removes the need for our hack.
hotfix/${version} needs to go into develop and master, so you need that tiny octopus merge to verify before starting the process
but yeah, incredibly rare.