Demystifying Git Submodules
cyberdemon.org
cyberdemon.org
Git is elegant is so many ways, but submodules are the broken, ugly stepchild of the beauty that git is.
I suspect it's not the idea of submodules, but the terrible, terrible command interface to them and how badly they work with the rest of the system.
I had created 2-3 char git submodule related aliases. Git submodule exists because a better alternative isn’t available (which gives us same behaviour or close to it).
More advanced things were created since then, like Fossil or Pijul. But the network effects make git predominant now.
That and the fact that it's... well, it's a good system - if perhaps not with the best UX.
It works well for pretty much everyone who cares to learn the basics, and then you can evolve from there with more practice. Which is probably true of any system.
But, say, Mercurial is also pretty good in many aspects. It used to be rather popular, but its popularity is waning, and not because of some kind of technical inferiority.
> or were too slow to for larger projects like Linux (Mercurial).
So it seems like there were technical reasons for Git vs. Mercurial. I don't really know and having never used Mercurial, I couldn't comment on how good it is or how it compares with Git.
From what I read around, it's mostly the UX that's marginally better on the Mercurial side. This is the point where the network effect certainly has weight. If one offering is not better enough than the one people are used to, there is no compelling reason to learn something anew and move all existing projects over.
When there is great technical reason, people will move though and the network effect will start moving across. See the previous systems that were popular before Git: Subversion, CVS, ClearCase, etc. Those have mostly been phased out completely, except maybe for older projects that have them ingrained into their processes and technology.
There's no wonder GitHub drew people in. It's interface and ease of building a community around a repo made of so much more accessible than anything else around at the time.
some suggested package manager, which imho is another layer of dependency, e.g. in Python, it always takes me some extra seconds to figure out which package manager I'm supposed to use, and it's constantly evolving.
end of day it's up to which one you are most familiar with.
I suspect the opposition to submodules comes from the poor (manual) integration.
As someone suggested in another thread, they ought to auto-recurse by default.
But I don't think that's going to happen, either.
Maybe something like direnv where it tells you about the unloaded submodule once you enter the repository. Except that won't work for IDEs.
I use them in my personal projects and have a sprawling mass of crazy. It "works" but it's a pain and I have scripts to recursively crawl through and basically do 'git pull' everywhere or 'git push'.
So if you end up using submodules when you don't really care about versioning the sub-repos, I think you end up feeling a bit of an idiot like me. But it's a pain to reverse.
It's just a rare feature so people don't know how to deal with it, so they just hate it. Once you figure out the 5 commands you will need you're golden.
At some point someone in a submodule repo is going to go rogue and make something incompatible with the bigger picture. Psychologically, they’ve merged to their trunk so as far as they are concerned their change is complete. Task done, “mission accomplished”. Yet until it’s also in the parent repo’s trunk, it might as well still be on a branch.
In a monorepo, the rogue coder would instead still be on a branch. They’re done when their branch works and their change is merged (er, ff’d!) onto the one true main, not before.
None of this applies to FOSS projects. Those are Big Tents where political negotiations are constantly required to keep various democratically equal projects aligned.
In the corporate world though I am entirely happy with a one party state running a top down, planned economy of N year plans following CEO-Thought, and with a single repository.
Pick your poison - change management is not an easy topic.
On top of that they are needlessly confusing. Why is there a .gitmodules file and hidden state inside the .git directory? Why aren't they cloned by default? Many of the UX issues have only been fixed if you turn non-default options on (e.g. the display of diffs can be changed from useless "submodules changed from commit 123 to 456" to "these commits have been added/removed").
Just all-round they are a mediocre idea, implemented badly.
Are there any downsides to completely skip dependency managers of specific languages and just use submodules to handle dependencies?
I don't mean code by 3rd parties. I mean the projects and libraries I write myself. I am tempted to try and handle my own code-reuse purely through git submodules. Would I encounter any problems?
Submodules suck once you encounter a diamond dependency problem.
Git's data system has 3 (& 1/2) types of objects:
1. A blob, which is analogous to a file, and is referenced by the hash of its contents
2. A tree, which is analogous to a directory, which contains blobs and other trees, and maps them to names. It is referenced by the hash of its contents.
3. A commit, which contains One (1) tree (the top-level of your repo), a reference to one (or more, for a merge) parent commit, and miscellaneous metadata like the author and the commit message. It is referenced by, you guessed it, the hash of its contents! (annotated tags are commit objects)
3.5. References, which are analogous to symlinks to commits. Branches and lightweight (non-annotated) tags are References.
Now, remember how a tree can contain a blobs or other trees? What if (gasp) you put a commit object in them!? That's essentially what a submodule is.
That's why a submodule is always included _at a particular commit of it_. That's why there's all sorts of complicated support machinery to make "a commit object inside a tree object" make sense.
Actually you can make a submodule track a branch instead of a specific commit. I've never seen anyone actually do that though and it seems like a bad idea. Though I did work for a while for a company that had written a custom tool that worked like that and we never ran into any problems due to it.
I think you're right actually the submodules. You can associate a submodule with a specific branch, but it still records the hash like normal and you still have to manually update it.
One project that I wrote, used nested submodules. There was a specific reason for the nesting, as it was a layered system, and each layer had a very specific context and functional domain, and submodules helped enforce that.
The problem was, it made changes a huge pain. If I made changes in the deeper layers, I’d need to propagate the changes throughout the entire chain, above. I wrote a few bash scripts to handle that, but it was fairly kludgy, and quite brittle.
I ended up just folding it all into a monorepo.
The one feature that I’d love to see in git, is something that Microsoft SourceSafe could do. I call it “Virtual Repos.”
You could make a “repo” that was actually an amalgam of files that were references to files in other repos. Their state in the virtual repo reflected their state in their “home” repo, and changes made in the virtual repo would go out to their home repo.
It must be a nightmare to get right, though. I can see why it would not be implemented, but these could be used for a lot of the same things that submodules are used for.
I don’t think sparse checkouts let you “mix and match,” the way virtual repos did.
Think of the files in the virtual repo as “symlinks” to their originals, in other repos.
Other than that, SourceSafe was kind of a nightmare, and I don’t miss it.
Most engineers have a poor understanding of Git. My university had a great history of version control course right at the dawn of the git era (In 2006! RIT really speedran it, standardizing on Git by the end of the year, but also including tutorials on RCS, CVS, and SVN and a brief foray into Perforce). Still, a ton of my classmates just didn't get it.
The other major blocker is what to do with an unclean submodule repo; I honestly don't remember what git does by default, because it's bad. And most projects get unclean real quick. Makefile hygeine is not common, and for most of time most projects became unclean from a simple `make`. It's better now, but not great.
If you're checking generated files into git, submodules aren't the problem imo.
I can see the frustration of modifying submodules files and trying to commit the main repo, but if you have to do that then it wasn't supposed to be a submodule. That's like complaining that modifying node_modules files doesn't apply upstream to your dependencies.
git update submodule --init
git init --submodules --update
git init --recursive --submodule --update
etc.Git checkout should do that automatically. Checkout a repo, all the submodules automatically checkout. There is no good reason not to.
Git reminds me of C - important language, used everywhere, but has never evolved with respect to usability.
EDIT: hmm... actually "git switch" was more usable to me
It has a different set of trade offs and works without any problems or changes to your workflow if they fit. (Only thing it has problems is rebasing, under specific circumstances)
- You can mix subrepo commits and main repo commits freely in a single commit, it’ll take care of submitting only the relevant changes when pushing upstream.
- Publishing changes from a subrepo iş just a single command.
- Subrepo adds a .gitrepo file to the subdirectory for metadata
The readme on repo does a good job of explaining things.
I'm sure there are advantages to git subrepo, but I am still not sure what they are.
For me, I just use git normally as I would, and do a subrepo push when it’s time. Subtree would make me change my workflow which is a big one for me.
There was not a week that went by without me having to unstick some team that had horribly managed to screw things up because of them or watch an engineer burn an entire day fighting with them or watch a new hire completely confused.
"Well if everyone would just ..."
Everyone is not going "to just". If your system relies on everyone inherently having the same understanding of the world and behaving in the same way, then it's a terrible system.
Edit: typo'ed command
[cjfd@cjfdpc ~]$ man 7 gitmodules
No manual entry for gitmodules in section 7Git subtrees also let you move back and forth between multiple repos vs mono-repo without losing history. So you don't need to solve that particular debate in your team.
This is where I use them. I have some Rust bindings to C++ code, and that C++ code lives in my repo as a submodule. Everyone seems to hate submodules I guess because of the surprising behavior described in the post, but for my use case they've been completely fine.
If you go back and watch Linus’ talk at Google regarding git, he’s basically describing (unknowingly) why Google needs to not use git for its day to day. Even on a smaller scale, Android (AOSP) had to create a meta tool for git called git-repo to handle its source tree. Git submodule failed there.
Where did you get that from? Sources?
Google rolled their own VCS, because Google is older than Git, and they needed something that works. Their custom VCS is a hacked up version of Perforce.
By the time Git came around, Google was already pretty much committed to their in-house custom tool, too many things relied on it.
The main thing that has been developed in git to allow very large repos is shallow clones (both in terms of history and slices of the repo). This model works well enough within git's logic, but it's just historically not been focused on until fairly recently (and I don't really know what the state of play is there - I think there's still a limit at a certain scale where simply finding the state of play of a large checkout becomes a bottleneck, and you start to want a persistent daemon to use FS notifications to keep track of what's changed instead of stat()ing every file in the tree)
(I've often pondered if it would be possible to make a DVCS where there's no firm repo boundary at all, i.e. you could construct a checkout from any combination of trees and commits stored in different locations, and have it work seamlessly. There's probably more than a few thorny issues in there, but it would be an interesting concept)
I recently did an embedded Linux design that depended on 5 external repositories: one from Yocto, three from OpenEmbedded, and one from a CPU vendor. My own code just sat on the top of this set of repos.
Submodules made that design very simple. One repo with all of my code in it, and submodules for all external repos. All dependent repos were pinned before of how submodules work. Pins were easy to update when desired, and never move on their own.
Yes, you can also solve this with a monorepo...
Almost every use case for a package manager is better served by Git, whether you choose to use submodules or not. If you want to do version control, use the version control system, and stop trying to do an end-run around the way it works.
Previously:
> I'm happy to criticize NPM the tool. The whole thing is designed as a second, crummier version control system that lives in disharmony with and on top of your base-level version control system (so it can subvert it). It's a terrible design.
Nope, Git is pretty good about downloading stuff over the network, too. In fact, it's so good at it that many people using a language package manager insist you use Git at some point even when (before) using the package managers. Indeed, there's been a lot of trepidation and gnashing of teeth about whether the places where language package managers download packages from are as reliable/trustworthy as the server where the Git repo for the software project is hosted.
> nothing about eg semantic versioning, or how to resolve different requirements from different libraries
"[…] incompatible reimplementation of _half_ of Git."
Remember, git was designed and written for the kernel first and foremost.
Note: I think GIT UX is horrible and requires multiple years of practice to be comfortable with.
Even still, submodules are also best avoided.