Show HN: Get rid of Git submodules and never look back (now for GitHub users)
gitmodules.com
gitmodules.com
Submodules are usually good for frameworks and libraries that are more stable than your main codebase. However, those submodules under active development within the same organizations can be a pain to maintain.
I'm not sure there's a good solution to this under git at all, but at least there's a product that makes an attempt and I absolutely don't mind the commercial intent here. For example GitHub is commercial after all, what's wrong with that?
Except they aren't. Libraries are better in their own code base and published as a separate artifact which is auto-versioned.
> I know some programming ecosystems, such as C++ does not have a default way of handling libraries and dependencies, and in that case it might be good to look into submodules.
> Please, if there is any dependency management in your chosen platform use it.
I use C++. There is no such thing as package management in C++ (at least no unified standard that you can use with confidence with every library). So giving a blanket statement and saying that you should just ship library artifacts is a non starter in a language like C++. This basically implies either creating a package manager, or distributing binaries manually, or forcing users to find the correct versions manually and install them.
Git submodules is extremely helpful here, and if you look at the documentation this is what it sounds like it was created for:
> Large projects are often composed of smaller, self-contained modules. For example, an embedded Linux distribution’s source tree would include every piece of software in the distribution with some local modifications; a movie player might need to build against a specific, known-working version of a decompression library; several independent programs might all share the same build scripts.[0]
If this suggestion here doesn't sound like your use case, then your probably using the wrong tool.
In the same way, you wouldn't tell people who use a hammer to cut a piece of wood in half that the hammer must be a terrible tool. You would say something like, you're using the wrong tool, here's a saw. Use the correct tool for the job at hand. If you need source code from a 3rd party library that you won't modify but you depend on, a submodule is a fine tool for the job.
A few big differences from submodules:
With X-modules, you have all the code in your repo, vs. a reference to the other repo. If you are looking at submodules to manage repo size, this solution is not for you.
With X-modules, you are tracking a branch, not a revision. If you want to pin an external dependency, it's not clear that you can do that with an X-module (the screencast doesn't show it, I don't think).
X-modules require an extra application to manage the repo relationships. There's a command-line client, but it's listed as unsupported on their website. It's not clear what the auth story is around who can make these changes. It doesn't look like adding an X-module would be subject to code review.
Personally, I'm not interested in this at all. I don't think it's a bad product, but I wonder what problems it actually solves. I don't know how common the submodule pattern is to compose a monorepo like they've demonstrated, not in an 'enterprise' setting they hint at future pricing for, anyway. And it's not clear to me whether their sync pattern is going to be OK for external third-party dependencies, unless you want to actually track their changes automatically vs. being intentional about when you pull in changes; this seems like a supply chain issue to me.
Benefits: easy cross-component changes, repeatable builds Issues: huge repositories, lack of code ownership boundaries, complicated versioning for individual products/components
Of course people nowadays rarely use submodules to build a monorepo, but I believe they would like to if submodules would be transparent.
Git X-modules are at early stage and only syncs a ref with a directory now. We'll add an option to pin a particular commit and most interesting - an option to receive module changes as PRs queue that could be individually accepted (and applied alongside with the necessary call site refactorings).
Git X-modules is available for GitHub users now, and we plan to add GitLab and Bitbucket support - both cloud and on-premises. So, we reuse platform's user management to only let repository admins/devs set up the modules - currently via web UI, and via .xmodules files in the future versions.
- you require that the submodule always move forward (reverts are not allowed; you must fix the submodule and get something there first) - the submodule commit must always be reachable from the configured branch (`submodule.NAME.branch` in `.gitmodules`) - you require that commit be in the first parent history of the branch
This makes conflict resolution trivial (the last prevents MR heads conflicting because the first-parent history linearizes and requires that "A contains B" be true in one pairing).
Without these requirements, it is really a wild west of insanity IME. Much better just to use the dependency mechanisms of whatever tools you use (assuming they exist).
Git has been adding more config settings to try to make it less painful, such as submodule.recurse, but that slows down most git-command significantly and the truth is most repos don't actually need to recurse - they just need the top level. So instead people end up writing aliases to try to handle it, which works ok, but it's brittle and people have to know to do it.
And if you have your repo on GitHub, they don't do some simple things that could have made it easier. For example, they could let the repo owner make the git-clone copy-to-clipboard thing have --recurse-submodules in the command you copy, so that cloning will do the right thing. But they don't. I understand why git itself doesn't want to make such things automatic, but GitHub could do it for at least private enterprise account repos.
BTW, I am _not_ advocating whatever this "Git X-Modules" thing is that the link points to. I have no idea what that is, never used it, and don't plan to.
They're still pretty annoying to actively work in.
We keep only DTOs in sub-modules so you really update it if you change property on some communication objects.
This way you also don't want them to be invisible - you really want to see and have people notice this was updated and actively choose what to do with it.
Say you're git, and someone has a submodule that goes from having:
sub/
.gitignore = a/ignored.txt
a/
tracked.txt
ignored.txt
to this, when they `git checkout` some other newer or older or totally-separate-history branch (possibly several layers above this submodule): sub/
b/
tracked.txt
what do you do with sub/a/ignored.txt? What if the whole submodule was removed?Git's answer to every single question like this is to either fail to perform the submodule version change, or leave a now-untracked-and-not-ignored sub/a/ignored.txt file in the directory, which leaves you with an un-clean checkout that it'll warn and complain and possibly conflict on.
It's highly unergonomic, but reliably avoids silently doing anything that might cause you to lose data. The awful ergonomics are what people generally hate about it. They'd rather it made some consistent decisions so it Just Worked™ in nearly all cases. But that's not Git-like.
https://github.com/ingydotnet/git-subrepo
git subrepos work simply by copying your dependency to a subdirectory and committing the changes using one large commit that retains metadata about the update to the subrepo. For that reason, git subrepos aren't symlinks. You don't need to git clone --recursive like with git submodules, and you don't need cross-repo authentication. Updating a subrepo means performing another commit.
Even though git subrepos are the most poorly maintained, the design is simpler.
I wish someone would fork and take over maintenance.
git subrepos are the best.
- you're using submodules to vendor dependencies or develop multiple related projects
- you have control over your environments/toolchains
You can generally use Nix to compose the projects together from separate source repos. You can have a version-controlled specification of exactly which commits you're depending on and control of when to update them.
Fair warning: I'm speaking with the naivety of someone who occasionally needs to build projects that use submodules, and as someone who has been tempted to use submodules for my own inter-related projects (but used Nix to avoid doing so), and not as someone firsthand familiar with other use-cases like working in a bigcorp monorepo :)
(On that note, are there any actual reasons to not just --recurse-submodules by default?)
All git repos end up in that folder - it's basically my project folder. Before explaining more, I should specify that this setup works well for me, but I could see how others may not agree with it.
Let's say I have 2 projects called prj1 and prj2. In that case, I would have the following directories:
/home/eric/repos/prj1 /home/eric/repos/prj2
Both of those would be git repos. Now, let's say that both of those repos rely on a shared repo, which I'll called shared-repo. In that case, I would have this:
/home/eric/repos/shared-repo
Then (and here's the part that probably won't sit well with folks), I just reference the parent directory from either prj1 or prj2 to access shared-repo.
Down sides: I'm forcing a hard coded directory hierarchy and directory naming. Up side: it's simple. If symlinks worked just as well on Windows as they do on Linux, I could just create a symlink in both prj1 and prj2 that points to my shared-repo directory and that could solve this whole mess too.
Having said all that, submodules have worked for me in the past, but I use them so rarely that I always need to look them up - so I've gotten in the habit of keeping things simple for myself and just doing things the way I mentioned above.
The biggest downsides of this approach that I discovered are that (A) projects are not self-contained, which meant some expert (i.e. me) always had to go and fix the deployments, because nobody else understood the system, and (B) a lot of the frameworks we used had their own system for organizing dependencies, which meant that we'd end up with projects inside projects inside projects anyway, and this system, whose intended purpose was to _flatten_ the directory tree, effectively just added another level to it.
Edit: (C), which is a variation of (B): if some projects are in monorepos, that throws off the system, too. Now instead of workspace/project you have workspace/monorepo-project/subproject, so now in some cases you have to reference "../../something" instead of just "../something". Not everything fits into the nice flattened-out system where relative paths are easy to guess, so you end up having to either merge things together with symlinks, which gets confusing because now you have two parallel structures going on, or just configuring dependencies with environment variables after all (at which point you wish you had just formalized the environment variables from the beginning to make everything explicit).
I'm not saying you shouldn't, as it works for you, and that's of course fine.
But, the key functionality of submodules is that the sumodule can exist on its own, as a git repo. The degree of dependency is then entirely controlled by the parent git repo.
which goes at more-or-less the same problem from the other direction. Josh helps with treating a monorepo as separate repos, as opposed to git x-modules' approach of treating separate repos as one repo. Either way, the idea is to try and get the benefits of both monorepo and multi-repo styles, while avoiding/mitigating the disadvantages.
HN talked about it 11 months ago: https://news.ycombinator.com/item?id=27844363
Anyone have experience/updates to share since then, or a comparison with git x-modules?
In either case, I'd love to see a comparison to both git subtree and submodules, on a more "workflow-by-workflow" basis.
This keeping both sides always in sync with each other can avoid some of the downsides of git-subtree, but has its own limitations like only working when both sides are owned by the same org, does not allow a subtree that maintains "local patches", etc.
If simultaneous pushes to super-repo and child repo occur at the same time, one gets applied and the other gets converted to a pull request (unclear if the server reports error back to the git client, or if the push appears to succeed, but then appears to be force pushed over by the other change, being preserved as a pull request.)
Maybe also I could set their dirs to be 'read only'; git should refuse any local changes and only accept a direct change of checkout (of hash or tag).
Every time I've encountered submodules in any work projects, I've personally replaced them with auto-incremented semver libraries (usually jars) and received copious thanks for it.
Submodules aren't ideal but I do prefer that you can pull down everything you need in a one liner instead of multi-stage pull, build, auth, download, build etc.
So I wouldn't say its better in every respect.
In my current job, which uses semver and conan and not submodules, submodules would have solved most of the problems. But unfortunately they are all git newbies (switching from svn just lately), so you cannot throw submodules at them yet.
Ask HN: Why are Git submodules so bad? - https://news.ycombinator.com/item?id=31792303 - June 2022 (167 comments)
Click-bait marketing seriously?