How Kubernetes Broke Git
matt-rickard.com
matt-rickard.com
If you've ever tried to develop software that depends on k8s modules you know what I mean - you inevitably get a diamond dependency conflict that go mod can't easily handle because some package needs version 0.45 of apimachinery but something else needs 0.46 (made up versions but you get the point). If they wanted to have many small repos they should have some rigor around versioning and public interfaces between those repos, rather than this magic manifest of specific releases that work together.
I know I’m being very uncharitable but I had to giggle at the irony here
This isn't any different than how say a b+tree (or other persistent data structures) rewrite their nodes from the leaf to the root, but leave non-involved subtrees as they were.
There winds up being a lot of activity on the superproject that amounts to just updating its submodules, but the commit log for it becomes a linearized history of stable/compatible commit versions.
There's definitely room for improvement wrt usability, but the claim that git has No atomicity across subprojects doesn't ring true to me
[1] or rather ../.git/modules/"$submodulename"/hooks/pre-commit, depending on how the submodule was added
> Or you could put a script at .git/hooks/pre-push.sample in the parent repo (untested) that verifies if all commits in submodules are in the respective upstreams
(a) those scripts are most certainly not called .sample (b) it was fast enough to set up a local test case and (as expected) each repo (outer and "inner" submodule) carry their own git hook setups and `echo "exit 1" > .git/hooks/pre-commit` does stop top level commits but does nothing for the inner repos
Yes, I carelessly copy-pasted the path.
And yes, as I said it would probably be a good idea to have a feature like that in the core. -- Ah, I see the other reply, so it's done. I generally recommend asking such questions on #git on IRC, someone will know the answer if a feature already exists.
git config push.recurseSubmodules check
Or you can make it push automatically by replacing "check" with "on-demand".You may also find it helpful to make various commands automatically apply to submodules with:
git config submodule.recurse trueI got excited about the submodule.recurse=true one, but at least for "git status" it did not descend into the submodule the same way that "git submodule foreach git status" does
I haven't experienced that scenario before but it seems like there'd be an obvious git error?
git clone --recursive some/repo.git
cd repo/subrepo
sed -i"" s/hello/goodbye/ README.md
git commit -am 'lololo'
cd ..
git commit -am'subrepo with *local* sha reference'
git status # everything is clean!
now that I know that can happen, running `git submodule foreach git status` will surface the "your branch is ahead of" magic text that indicates what has gone on, but it would be tons better if the system understood what was happening and didn't allow such a bad outcomeI'd hope so: the Linux codebase is an order of magnitude bigger than the K8s one and it's not breaking Git.
The model of the “benevolent dictator” kinda works in this case. My last project I was on we managed to scale to about 50k LOC without anything special, the key is I knew the repo like the back of my hand and could catch potential integration issues. While the model works well, it’s very hard to setup as you need a real nerd of a team lead to constantly watch the repo.
Most things in life that don’t have leadership become messy and disorganized and eventually disintegrate into an unstable hell with no real focus
If you misuse a tool, and the tool performs poorly at the job it's not designed for, but never fails in an unexpected way, and still maintains all the functionality it always had, how have you broken that tool?
If try to hammer in a nail with the butt of a screwdriver, and make a complete pigs ear of it, but the screwdriver absorbs the abuse and is still perfectly usable as a screwdriver afterwards, did I "break" the screwdriver?
Or, am I misunderstanding how the word "breaking" is being used here? Is there a meaning I'm not getting?
> authorization
> package management
> So why shouldn't a VCS embrace its role as a collaboration tool and explore more generic merge-based optimizations like a queue?
In the bottom section there are a few 'wishlist' items, can I call it that. But those aren't good VCS features, they're a reflection of the k8s development world which is not how most of us do development.
It's also assuming that because k8s is an all-in-one-doing-many-things, that the VCS it uses should also be a huge all-in-one. I don't think it should; all that would happen is the leaking of k8s' already complex existence from k8s into git.
Then it really would break git by making git worse for everyone. I would suggest finding another tool, or making your own.
https://josh-project.github.io/josh/
It basically lets you unify/view many repositories as a single one, or equivalent to split a mono-repo into smaller sized units of work for CI, specific teams, etc. It's bidirectional, so you push and pull from josh and everything goes into a single linear history in the mono repo. And because it's bidirectional, people in the mono-repo can still do things like make large-scale atomic changes across all sub-repositories, and those get reflected.
Josh currently isn't suitable for a lot of workloads due to various reasons (authentication is one that stands out), but it's actually the first tool I have seen that manages to offer BitKeeper-like "subtrees" that work really well, at scale, for large repos and teams. It requires some care to make sure "sub-trees" can be usable units of work, but it was one of the best features of BK in my opinion and really great for people doing one-off contributions, or isolating trees/changes to specific developers.
I'd be interested to know if there are other open alternatives to this. It's a nice point in the design space between solutions like "integrate with the filesystem layer to do sparse clones" or "just split up the repos."
In fact I think authenticating and authorizing access to components of a monorepo is definitely in scope for its design, and could allow really powerful and cool things. But I just don't think it does that yet. Maybe it shouldn't. That's all I mean. It's not like a hard rule. There were some other things that I thought might hold me up, but I can't remember them now...
Not sure if you've ever used it but there's some self-hosted software out there called 'gitolite' that does a lot of this. Perhaps they can be made to work together seamlessly...
That being said, the overall scope of permissions and authentication is more complicated, more improvements is definitely needed
There is already an idea and initial implementation of path based ACLs for Josh. However, even if that concept was perfect and already implemented, it would be kind of useless.
Why?
To be useful in practice we would need a UI for code review in a monorepo. This UI would need to aware of the "workspaces" and ACLs in the repo and respect them when showing files and diffs to the users.
As it stands now, Josh is used together with either Gerrit or GitHub (or similar). Patches or PRs are always being reviewed in the context of the full backing monorepo. As long as those are the only options to do code review, I don't see the value of having access control at the Git level.
That being said, I am planing to create a new code review tool that does support these things, but it has a long way to go before it will be a serious alternative to the common tools used today.
Choosing a monorepo vs many-repos is a tradeoff, there are consequences to choosing a monorepo, many solvable, and a tool like this just provides robust solutions to a couple of them, is what I'm saying.
Actually one very common use case Josh solves that I've had in the past is "Merge a repository into another, while developers keep using the original as if nothing happens." This is important to keep teams moving while a migration happens. I have this problem right now at work; two repos that want to be one repo, one smaller and one bigger. You can use git filter-history to do this but Josh is significantly more powerful, and most importantly the team whose repository got "merged away" (i.e. got merged into the bigger repository) can keep working on their repository as if nothing happened, and you can eventually switch them over to the main repo. Normally you have to stop the whole train at once and move people over while some poor bastard has to surgically modify the git repository after they take the git repository down. But Josh allows to you to merge and incrementally migrate that repository, because you can now view one repository as a "workspace" of another. It's sort of like the difference between taking an optimistic lock vs a normal wait lock, in my mind. Josh lets you do "optimistic locking" when merging two repositories, rather than making every team stop -- serializing -- while it happens.
Another common case is "I need to mirror a subset of my proprietary repository onto GitHub." Actually "I need to mirror subset X" in general. This is another monorepo problem that, while not super complex, is actually really nice to solve in this way because bidirectionality means people who patch the mirror downstream can still have their changes merged upstream. This isn't always possible for QA/workflow purposes (e.g. their downstream change could break something upstream, so many people choose to instead apply it the other way around), but it's something I've experimented with. Not unthinkable.
I do think it's a very promising project with many real world applications. I'm still figuring out how best to organize and use it, just like we do with git.
The merge issues seem like they would be solved by your code-hosting platform. (GitLab has Code Owners and Merge Trains and I imagine GitHub has something similar) To me, these features are something you'd implement in your centralized tool rather than git which has to support a decentralized workflow. Perhaps someone clever could think up a decentralized authorization system for git, but is it worth it when almost every project has a centralized source-of-truth repo?
> A system that could record atomic commits across projects or a better submodule experience would have allowed for more flexible developer organization, especially as the project grew to a new scale.
but otherwise I'm with you that this could have used a better title or something
I see the latter much more often when I jump companies and Google’s projects have had terrible API stability so I’m not really sure Git is to blame here
I can't stand not being able to run everything in the same window with ctrl P picking up files from across projects as a reference.
I feel like I'm the odd one out because I've noticed a lot of languages and Lang servers are making these assumptions about how devs work and organise code.
Or they're just being perfectionist opinionated twats.
In modeling & simulation, this is called "emergent behavior". While that may be imprecise in terms of the definition, stand by for the effects.
Doing anything at scale separates the pros from the dilettantes, e.g., me.
The way they are operated just doesn't fit the way any human thinks.
What I wanted to do was be able to have people work against a pinned version of a different Git repo, then update the sub module whenever we felt the need and handle the build breaks. This task seemed impossible to do correctly over time, which I just could not understand. How was the submodule getting updated when I didn’t call anything? Why are submodule changes appearing in other commits? I just couldn’t figure it out.
I am joining a new project and they started talking about submodules and I what I said was “yeah uhhuh cool” but inside I was pretty nervous . But I couldn’t be sure it wasn’t because I was a dummy and they knew exactly what they were doing, so I kept quiet.
So just to be clear, two of the most fundamental day-to-day operations you can perform are turned into massive liabilities from this feature, ones that are likely to either break your build and/or just make your life harder. As someone who had to maintain stable and development branches of a project, cherry pick between the two, cut releases, etc, submodules are simply hell, because they make an already difficult job worse. This is a good sign that they are a liability. In a past life we actually had so many people push invalid submodule updates over time that we eventually wrote a git hook on our server to reject all commits with submodule updates that didn't exist in the corresponding repository, and that were not specifically tagged in the commit message as updating a submodule (through a magic set of keywords.) The fact we even had to do this is its own pain.
I have maintained projects that have long-standing histories with dozens of submodules. And every single time we eliminated one of those submodules (often by merging into the parent repository, or simply dropping the dependency entirely), we all breathed a sigh of relief, and our lives all got significantly better from that point forward.
As you can tell, this experience has made me very prepared to fight against submodules everywhere I might see or encounter them. But trust me: it's for your sake, not mine; 'cause there ain't a chance in hell anyone is adding any to my repositories.
I will admit that if your dependency in this case is something that changes extremely rarely, and will only see updates maybe like, bi-annually, submodules are "OK." Not great, but they'll work, and presumably won't inflict massive psychic damage on your team members. But it's once they receive any foot traffic by more-than-one-person that all the real pain begins.
Our product has a somewhat simplistic git interface (behind the scenes it's anything but) and I've tried to keep it so, however lately customers have started demanding we also support submodules.
The problem is that we use git to mirror a hierarchical database, so using submodules means mirroring another hierarchy inside our hierarchy. This would mean changing the current assumptions in the code to ignore things in the sub-hierarchy except for the sub-sub-hierarchy we care about. Yeah this is hand-wavy but the design constraints I've had are kinda hard to explain.
Also, sub-modules would require changing all our git calls to take submodules into account, including cloning, reset, branches etc.
And I've read many people's bad experience with sub-modules, including the ones in this sub-thread and so now I'm afraid they might hurt our maintainability in the long run.
I've thought of a couple of ways to do this without sub-modules, including using git subtree, but all of them have drawbacks.
I've actually found a neat way to merge a sub-directory of another remote repo to the current repo - which means we wouldn't have to change any of the existing code. It involves only "standard" git commands, basically only checkout, reset, and merge (without "exotic" commands like subtree, read-tree, and whatnot). And using "exclude" to keep only the content of a single subdirectory of the remote. But it does require us to maintain a file that's exactly like .gitsubmodules to keep track of remotes. And that's the thing which git submodules does for us "for free".
Also I've developed a bad state for bespoke solutions and NIH. I already fear I have contributed more than enough NIH to my company by developing the existing solution, but given the conditions I think it was the only logical solution (a previous bespoke solution failed and was cancelled).
But the longer I read and experiment with submodules it looks like they are also a kind of bespoke solution around basic git, and essentially require changing the way you handle operations such as reset, checkout, etc. Training all our users to fix errors due to out-of-sync submodules looks like a nightmare, when they already have problems with the current solution and with git in general. So I'm really conflicted on what our current path should be.
But the description of this use case / requirement sounds complicated / confusing to me ("we use git to mirror a hierarchical database, so using submodules means mirroring another hierarchy inside our hierarchy").
I don't know what your product is but I hope for your sanity a PoC is possible.
An obvious solution is modularization and stable internal ABIs, but the Linux community have avoided that approach for a long time, and with good reason.
IIRC he's already the maintainer for the "stable" branch so a lot of the work people think Linus does is already on his plate.