Don't couple your Go code to GitHub
iain.rocks
iain.rocks
I’d remove “go” from the above, i.e. I think same applies to other stacks.
Even using GitHub domain links in code comments gets problematic long term. Ie when a migration happens and those links start pointing nowhere.
Why can't you just search-and-replace? Presumably all of them refer to "GitHub.com" and not much else code will, so I'd think this was an exceptionally easy case.
Even easier for comments, since them being obsolete for a few hours during a migration doesn't exactly break anything.
It was horrible. A team of people spent weeks. Different projects depended on different versions of the same internal libraries, we had to create a branch for each depended-on commit and make a version of that commit with the new URLs. The expressed goal was to end up with a system that's "exactly the same" as the old, just on a new host. Changing which version of libraries projects depends on would introduce unnecessary risk.
And the result is a code base where bisects are broken and where it's impossible to build an old version of any of the code without a ton of work.
This is one of the main motivations for using a monorepo with all third party dependencies imported all the way in, and only one version for everything.
Of course it causes a whole bunch of other problems and is kinda expensive to scale, so most places won't do it.
Pick your poison. Each approach has some seriously negative trade offs in the extremes.
That makes no sense? As long as you’re using go modules, you can just host an internal go proxy for internal modules similar to proxy.golang.org to archive the old versions, and they won’t depend on the old git host.
I say this because I’ve gone through a few of these, some with monorepos (the easiest, search and replace and you’re done), some not (yes the hardest but you can just vendor).
At least in the situations I’ve faced this was genuinely not horrible beyond just the organization itself being complicated about the solution.
Clone Server A -> Server B. All code still refers to A.
Update code to refer to B
You can now EOL Server A, and B becomes the new canonical reference.
I think the key is simply that you can keep both A + B running during the migration, so you just need to be able to do a code freeze for the duration of the migration. And a single person can easily do this migration with a few python scripts and an hour.
Bonus points: code freeze guarantees your #2 - no changes to how the code works.
Of course, I'd assume that most complaints come from situations where this "secret trick" isn't viable
The "module name is network path" is a convenient convention but not at all some "limitation" of the tooling.
In Go, the package identifiers are literally URLs. Go's own tooling assumes that it can make an HTTP request to the URL and that the response will be HTML with particular tags which Go's tooling will parse and use to resolve a git repository which can be 'git clone'd. Leaving them as-is when you abandon the infrastructure they reference literally means leaving dead links in your source code.
Any tool that automatically assumes that a URL in a Go module import is 'trusted' is just a broken tool.
If you haven't added a replace directive to your go.mod, then you certainly haven't replaced all instances of the old URLs with updated ones.
Heads up for anyone who doesn't know: this only works at the "top level". Any replace directives in your dependencies will be ignored[1].
So for example if you have a dependency tree like [main -> thirdpartyframework -> golang.org/x/net/http2], and thirdpartyframework uses a vulnerable version of `golang.org/x/net/http2`, you can't just fix it by patching the thirdpartyframework repository with a replace directive; no, because that would be too convenient. Instead, the replace directive needs to be at the main module, where it doesn't make sense and is inconvenient.
Even though I like Go, it really seems like they don't care about anything other than monorepos. As soon as you need to work with forks, mirrors, or even just private modules[2][3], the tooling actively works against you. Also using your custom module proxy is a pain.
[1] See: https://go.dev/ref/mod#go-mod-file-replace:~:text=replace%20...
[2]: If you've only used private modules hosted on GitHub you might not have noticed too much pain because the Go tooling has hardcoded behavior specifically for GitHub and a few mainstream forges. You don't find out about this until you try to self-host something like Forgejo on your own domain thinking it would Just Work(tm), but it doesn't, and now you're left wondering why tf it works with GitHub but not with your own forge instance.
[3]: I think there's no hardcoded code for SourceHut, so you might be able to experience the inconvenience by hosting private modules in there.
Wait, what? If you decide, in your own application, to force all your dependencies to use a specific version of golang.org/x/net/http2, then obviously you'd want to be able to put this directive into your own application's source instead of going around patching 3rd-party repositories (that's just rude).
Even in the case where it’s an internal only Lu array, it can be complicated to refactor a library name.
I did not realize it ran that deep - I'm used to distributing EXE files, not anything that would need to know internal dependencies like that.
But it's something you have to do as an active choice; the tooling, and your colleagues, will nudge you towards using URLs to your primary git host's web front-end. And it naturally doesn't work for libraries.
Many tools will fail if you break this contract.
Everything else is fine. Generics, errors, whatever...
I have seen too many cases of requirements on internal projects that prohibit devs from working because oops the vpn is down now or oops gitlab is under load and the pipeline for your dependency won't finish for the next hour.
It's not only an issue of availability or depending on circumstances you can't control, but a matter of reproducibility hence dependability.
I want to be able to compile critical software when a zombie apocalypse hits, and I have a growing discomfort about trusting outside parties for their availability and benevolence.
Even if you don't want to go as far as vendoring (which can get a bit out of hand with, say, NPM), a middle ground is using Artifactory or whatever as a proxy.
Then be more deliberate about when you update. AI may even help here where static analysis doesn't, because you might be able to use inference to see if a dep upgrade is even needed. Unless it has a severe vuln you probably don't need to track latest.
In this terminology, the article basically amounts to saying "This namespace can fail! The solution is, use another!"
But that doesn't get you anywhere, because the new namespace can fail too. In fact it is almost certain that it will fail sooner than "github.com's DNS and hosting" as a namespace.
There is no solution where you tie your software release to a namespace that can't fail because there is no namespace that can't fail. It doesn't matter if your favorite language uses DNS or has a centrally blessed repository or if it distributes a blessed list of package names with the language itself or anything else. The namespace can fail.
Therefore, the only thing you can really do is be resilient against failures, and in a lot of ways, the only practical solution to resilience is just to assume that if, in the future, someone has problems getting a package due to a namespace failure, they will not helplessly disintegrate into a puddle of tears while your code is lost forever, but that the future person will instead solve this perfectly solvable problem.
> it will fail sooner than "github.com's DNS and hosting"
When it does fail you can fix it. When github fails all you can do is stare at their status page.
You can also just use "replace github.com/example/example => gitlab.com/example/example" in your go.mod file and everything will keep working. That seems like a very pre-mature optimization for something that doesn't really matter.
You know, normal development task, small amount of story points..
For my thoughts about why it's not easy, see this comment: https://news.ycombinator.com/item?id=49870363
Now you want to move from github.com to git.example.org. You change libfoo, authservice and apiservice to use git.example.org, that part is just simple tedious work. But the v1.2.0 and v1.3.0 tags of libfoo are old commits from before the move, so they still reference github.com! Now you need to branch off of the v1.2.0 and v1.3.0 tags of libfoo and do the same change there.
So we have only 3 repositories with only 2 dependencies and we already have to do the search/replace 5 times.
Imagine now that apiservice depends on authservice v2.3.7 just to include some type definitions. Authservice is currently on version 2.4.0 but the types haven't changed so apiservice hasn't upgraded its dependency. Now you need to branch off of authservice v2.3.7 too and do the job there. Oh and authservice v2.3.7 depends on libfoo v1.2.5, so now you need to make a branch off of libfoo v1.2.5 with the search/replace.
3 repositories with a straightforward dependency relationship, 7 search/replace jobs.
The numbers get terrifying as you scale this up.
The idea is: the package maintainer pushes a final version of the old package that is a shim, and which has a deprecation notice. The shim imports the new package, and forwards all calls to the new package (and I suppose type aliases and variable aliases as well). The shim includes `//go:fix inline` annotations on all exported symbols, so that when users run `go fix` it rewrites their code to use the new package, by inlining their usages of the old package (which is now simply a shim that references the new package).
Not perfect. There is still a bit of a discoverability problem. Not all tooling warns on deprecated packages/functions (go toolchain doesn't, but gopls and staticcheck do). And users need to know to run `go fix`.
1: https://go.dev/blog/inliner#example-renaming-ioutilreadfile
// Deprecated: use example.com/mod/v2 instead.
module example.com/mod
As library user, it's your responsibility to keep your stuff updated. There may also be utilities in the various automatic dependency updaters that can migrate these things."replace example => github.com/example/example" at first, then change to "replace example => gitlab.com/example/example" or similar when there's a migration
(i don't use go, but i hope this is possible and i am confused if it's not)
One day, we’re all going back to vendoring dependencies.
It bloats your repo, both with the actual code, and the large diffs when you update it.
You have to manually track new versions, without something to tell you if new versions are available, or if your version has known security vulnerabilities.
If the dependency has it's own dependencies, you have to vendor those too recursively. And if multiple dependencies have the same transitive dependency, it is up to you to deduplicate them, and make sure you have a version compatible with all dependents.
Etc.
Recursive dependencies have the same issues whether you vend them or not.
Ensuring that diffs to updated dependencies remain within a vendor folder is trivial.
So, what’s left?
The biggest problem isn't (usually) disk space, or network bandwidth, it is that git operations slow down as the size of the repo grows. And it means that cloning or pulling the repo takes longer, which can be especially problematic for CI.
> Recursive dependencies have the same issues whether you vend them or not.
Package managers usually handle resolving recursive/transitive dependencies for you. Some have support for vendoring dependencies, but not all do. In theory, you could have similar tooling for vendoring dependencies, but in practice that often isn't the case.
That sounds like it's inconvenient to me.
And both of those concerns can be addressed by using an internal/private registry that mirrors the packages you need.
But in any event, your original question was to elaborate on why vendoring is inconvenient. Whether or not the benefits are worth the inconvenience is a different question, to which IMHO the answer is "it depends". Sometimes it is, and sometimes it isn't.
How so? Whether a dependency is vendored or not, you still have to update it to integrate a security update to that dependency, don't you? The alternative is to not pin your dependencies, but that is far riskier overall.
> Whether or not the benefits are worth the inconvenience is a different question
I would contend that it is the most important question. :-)
If I want to upgrade the disk on my MacBook Pro, I need to buy a new MacBook Pro with a larger disk. If I want to upgrade the disk on my work laptop, I’m SOL.
> Ensuring that diffs to updated dependencies remain within a vendor folder is trivial.
It’s not obvious to me how putting the dependencies in a vendor folder solves the diff problem. Does every code host allow you to hide diffs to certain directories?
And what’s the advantageous scenario for vendored dependencies? Is it just when the mod proxy and the upstream code host go down at the same time?
Is this a significant risk in reality? MBPs today come with a minimum of 1TB of storage. Even 5 years ago I think the minimum was 256 GB. This is more than large enough for all but the most massive repositories, even with vendored dependencies. And you can always plug in external SSDs or HDDs or connect to a network server.
> And what’s the advantageous scenario for vendored dependencies? Is it just when the mod proxy and the upstream code host go down at the same time?
This is a useful homework assignment. Ask your favorite LLM or consult some respected release engineering books. Also consult your local AppSec and infosec teams.
Consider the possibility that the drive needs to accommodate more than just a single repository?
I have lots of software, entire language toolchains, AI models, Docker images, application volumes, VMs, etc.
And this isn’t a theoretical concern; I have to free up space every few months.
> This is a useful homework assignment. Ask your favorite LLM or consult some respected release engineering books. Also consult your local AppSec and infosec teams.
So there is no advantage scenario that you’re aware of?
I use dotnet and I never liked seeing dlls and binary files in my diffs. I would argue if we are adding vendor code to our projects, we should demand the FULL source code instead of dlls. Maybe it is already possible with things like x unit. I have never given it much thought... But then that vendoree code has to come from somewhere as well, right? I mean there is something to be said about provenance or something here?
Sorry if this feels like a stream of consciousness because it is ↔
The answer to the question of why THAT is the case - is not so easy to answer. The main benefit I can see with using lock files instead of vendoring is that it saves a lot of storage and diff history from entering your repository. So clones are much faster, backups smaller etc.
I think Go used to work this way (automated vendoring) but it’s the only language I can think of that ever did in terms of standard tooling. It would be instructive to learn why that changed.
Maybe at that point it won't be using git either, or maybe it will - who knows.
Signed modules / including the signature hash would also solve a lot, e.g. it'd mean domain sales no longer silently inherit full permissions. It's sorta a shame that Go keeps doing such a good job at a minimum-viable wheel-rewrite, but then lets it linger for so long without catching up to the rest of the programming world.
How many mainstream languages have content addressed imports? I can’t think of any, so I assume I’m misunderstanding your meaning of the term because you seem to be suggesting that it is common and Go is the outlier for lacking it?
Which keeps happening with stuff they rebuild from scratch - an excellent and somewhat unique first showing, far beyond what most first attempts manage, but followed by near-complete stagnation while issues that everyone familiar with the field predicted from miles away pile up.
Which is a shame because there is quite a lot to like about Go in practice. And in spite of it all I'm thrilled that it is eating into Python's share in a lot of places.
You already can use SHA in go.mod in exactly the same way you would use a version string.
I don’t know about using multiple proxies in go mod though.
And while the go.sum file in a module is returned by proxy.golang.org (somewhat surprisingly), that only includes the module's dependencies, not itself. So you're still stuck trusting a goproxy to serve you the correct data.
If github is "taken down" you can't just resolve the whole domain differently, you need to only resolve the packages differently. If your "Golang domain" is taken down, it should be a lot easier to hotfix until something proper is implemented (the quickest and dirtiest would be a hostfile-entry).
It is being retired because very few people use it. It kind of sucks but this is how domains work, it’s not real estate.
Deepcopy [1] is a Go package that is still heavily used despite its creator has disappeared 9 years ago. At least GitHub is a stable and trusted host as a distribution point and communication point for users.
It is bad advice to move your packages under your own domain. You will never be as good as Microsoft to keep paying for the domain. There are no guarantees in life but I do guarantee you that when you go out of business, that domain is the last thing you will think about.
If I own the domain, I'm likely to continue to owning it until I shut down my company. At which point, no one is likely to be running my code.
Possibly different if I was developing an open source library or something, but either way there's risk and need to be thoughtful of the risk versus impact.
It could be a URN that used the DNS as a back end (though that’s a lot like a URL) so better would be something more abstract with multiple possible resolvers and a signature.
Then the rest of the article explains why this is NOT a good feature in practice.
I don't think it's unworkable either, but this is one of these little thing that Go decided to do different and convinced its fans that this is a great idea and all the other languages where doing it wrong. After a couple of road bumps appeared, instead of admitting there are some advantages to having official package names, we're now told that everybody should just set up their own custom domain with an nginx server or a Go Vanity URLs forwarder to serve traffic for their GitHub-hosted packages.
The one thing that I was worried about was returning 301 in the example Nginx config. If you ever wanted to change the url that clients are redirected to then any browsers that visited the old url config would be forced to go to the old config's redirect url. For `go ...` and `curl` it wouldn't matter, but Chrome/Firefox will cache that 301 permanently and break the intended redirect. Not sure if this is really an issue in practice though.
You now have many choices on how to proceed, but none of them will include A) being able to build old releases, or B) doing so without making changes to all dependencies.
One choice is go to the leaves, C in our example, and make a release on gitlab. Then go to B, and have B depend on the gitlab C. This is fine for rolling forward, but if you want to roll back you would have to rewrite all of C to use the gitlab url, and find all of the equivalent tags and repush all of them. Also tell your users that any binaries you've released now have new hashes. Then repeat this for every repo. It's not a small task.
That said, in the normal run of things you can mark your last release on e.g. github as deprecated, leave a note that it's moved to xyz, and wait for the users to migrate themselves. Migrating is fairly trivial in that it's a find / replace.
However, Golang saves the hashes of all the modules in `go.sum` files, so we just need to add a way to do content-addressable fetches. This solves the issue of reproducibility for old builds. As long as you can find the module in some repository somewhere, you'll be able to build it.
The next step is supporting module _evolution_. We need a way to declare: "From this point onward, `github.com/company/someproject` is now `company.com/someproject`", so that all the references are to these packages are identical. This is possible on a per-package basis with `replace` directives, but this doesn't scale.
And this is not easy to solve in general (especially in the age of supply-chain attacks). If the initial project cooperates (or if Github can be convinced to help), perhaps at least a part of this can be solved by adding special "redirecting module" support to Go.
Unless of course, it's special domains that has identification built in, such as .onion which is generated in such way (cryptographic keys) no other people can easily obtain control even after the domain is no longer maintained.
If you really don't want to use GitHub, an alternative is just use .internal suffix (i.e. yourproject.internal/project) in combination with `replace` directives in go.mod. But that require your user to manually download/install your package and then edit their own go.mod.
The post is about commercial software. If enterprise’s domain is not credible enough, then there are bigger problems.
Well, I too dislikes their enterprise tune, "Changing is consistent, GitHub give you advantage (over other companies)" blahblahblah. Many developers I know are on GitHub because it was/is the place for individual developers to share code, not because the "company advantages".
HOWEVER, I can't ignore the fact that individual developers just can't pay enough to keep the machine running, so it's reasonable for GitHub to pivot towards commercial market. On the other side, GitHub is still benevolent enough to provide some enterprise-class service back to freetier users. This arrangement checks out for me so far.
I do self-host some of my projects too with Gitea. But I never dare to use my domain in my Go packages, due to long-term security concerns.
When I initialize my projects, I just use .internal domain, then if I decided to upload the project to GitHub, I'll change the URL to GitHub to make it more convenient to the users. Otherwise, manual download and `replace` directive like I suggested.
> After spending years doing those FFI dances in both directions, I’ve reached the conclusion that the only good boundary with Go is a network boundary.
It looks like this extends into the package manager too. The only good boundary with Go programs is a socket. The only good boundary with the Go package manager (? if we can say it is one) is a domain and a DNS server (your own or somebody else's in the case of gopkg.in).
[^1]: https://fasterthanli.me/articles/lies-we-tell-ourselves-to-k...
MVS + OCI is the gold standard imo
CUE module proposal (implemented a while back) - https://github.com/cue-lang/cue/discussions/2939
MVS, the fast and deterministic dependency resolving algorithm Go uses instead of SAT solvers - https://research.swtch.com/vgo-mvs
-- Curious how this avoids SBOM need?
Alternatively, could we ourselves build an automatic mirror so there's a redundant supply provider that doesn't depend on a maintainer's choice of git forge
Let's take a project I worked on for a couple of years some time ago, it has this go.mod file: https://gitlab.com/nunet/device-management-service/-/blob/ma...
Is it just changing all instances of `github.com` to `proxy.golang.org`? Shouldn't `go build` warn about using third party proxies for dependencies? Should `go get` offer a flag to automatically use the proxy when adding dependencies?
Here's mine: https://github.com/foundata/hugo-theme-govanity (e.g. used at https://golang.foundata.com/ )
And yes I'm aware of the irony of hosting this on GitHub... still figuring out a good workflow for maintaining our OSS on Codeberg and GitHub in parallel fed from the internal forge. The dependency on our own domain is real but we favor it.
Leave the Go code or Linux distro to link-rot for a couple decades and all imports like this will be dead or serving malware, leaving the software unusable.
"This namespace can fail! The solution is, use another!" while i agree with this but setting up different namespace just because you want to avoid vendor-lock , we need to think from maintainer's perspective as thi would introduces frictions and so are of the most open-source maintainers willing to take it or not?
Not to be cynical about this but I fail to see how running sed s///g after a very rare event qualifies as a serious problem.
Now, sure, given that the solution is so simple, it's a nice recommendation, but still...
We should lean towards IPNS (Name Seever part of IPFS) or some distribited ledger.
Compiling code shouldn't hit network by default.
This is something Nix and Guix are trying to achieve: by default, they fetch source code from the original endpoint, but if it doesn't respond anymore, they would query the Software Heritage archive [1].
Because go mod doesn't use SHA, it makes things a bit more complicated but Nix/Guix encountered the same issue and there are some mechanisms to workaround this issue.
[1] https://www.softwareheritage.org/2025/05/21/software-heritag...
Bonus points if DNS mafia deems your cool TLD to now cost 10x times more coz it's "premium" and you're stuck paying up or telling everyone to migrate
People can delete projects from GitHub. Businesses whose priorities might change may not keep an old project set to public up on GitHub because they won't want people to continue contacting them for support, or they don't want to be responsible for updating security vulnerabilities for projects they are abandoning so they'd rather just pull it offline, etc.
Whether the library you're importing is hosted on GitHub or not, never assume it will be there tomorrow. *Always vendor your dependencies.*