The GMP library's repository is under attack by a single GitHub user
gmplib.org
gmplib.org
The FFmpeg-Builds repo has a GitHub Actions Workflow which clones the Mercurial repo. However, this runs as a (daily?) cronjob. In addition, this repo has 700 forks and now all of them are running the same workflow.
This is out of control for the original author of the repo as there’s no way to change those 700 forks…
EDIT: Also relevant is that there’s no official GMP mirror at any of the major code hosting sites (GitHub, GitLab) etc.
> GMP is at this point the only dependency that does not offer a sane way to clone its repository.
The latest release is from years ago and https://gmplib.org/download/gmp/gmp-6.2.1.tar.lz should contain every file necessary. Why would you need access to the commit history in a build script?
In fact, why would you need to download a fresh copy of a dependency that almost never gets updated from upstream? Surely Github can do some kind of caching.
I wonder what will happen now that their build servers can't even reach the Mercurial repo anymore. I suppose the problem should quickly correct itself soon with all the failing builds.
The nature of this attack is that 700 forks with each their isolated CI containers is downloading the file oblivious to all the other downloads. If GitHub were to cache this download, they'd have to man-in-the-middle, but people are using curl/wget and not a package manager with caching mirrors.
I'm not even sure cleaning up these 700 DDoS'y forks is that easy.
It needed the commit history specifically because the latest release is from years ago.
It kinda makes sense, but not in the context of how most people use the "fork" functionality, which to me seems to mostly be a "create a backup of this repo". So GH is building this project 700 times in different forks, for people probably not having made any changes to their fork.
CI providers really should do more to prevent and mitigate these when they happen. They should have outbound firewalls, and the ability to request a rate limit on IPs. Having to resort to a ban hammer when people depend on your tool is a shame.
https://gitlab.com/gitlab-org/gitlab-runner/-/issues/991#not...
Reply from Github: https://gmplib.org/list-archives/gmp-devel/2023-June/006162....
What is the basis for this quote? Ctrl+F-ing on what look like the relevant pages turns up nothing. Maybe I missed it, but Google seems only to know about 1 instance of phrase from the last week (and it's the one in your comment).
In before “it’s just one guy”: there are a thousand ways to solve this that are not expensive or complicated.
Should the offending party have been a better consumer? Yes.
Is it fair to entirely blame the user versus implementing additional caches and safeguards to make it an annoyance and not the end of the world emergency GMP is making it out to be? Definitely not.
https://github.com/BtbN/FFmpeg-Builds
There's a bunch of build variants here, for various platforms, in different ways. Classic "you have to create the cartesian product of all variables" approach. Here is the single build matrix for a single "Build FFMPeg" job:
https://github.com/BtbN/FFmpeg-Builds/actions/runs/530351117...
You can look through the code and see that each of these will trigger a download.
The developer also added an interesting commit today, stopping forks from having the CI work instantly. Why? Because by default they would all share the same cron entry, and therefore launch all at the same time every night. So now new forks will need to go in and tweak/spread the Cron timer around if they want builds of their own
https://github.com/BtbN/FFmpeg-Builds/commit/78191a73a6a1959...
But there are already over 700 forks of this repository, so the current path is unsustainable.
So in short:
- This isn't an "attack", and it's unclear of the direct volume, but anyway, it doesn't matter because ultimately gmplib.org had strain caused by it.
- GitHub correctly identified the responsible user when asked, and it did not in any way seem malicious, but certainly excessive.
- The build/CI infrastructure for this project definitely hammers the project more than it probably should, exact specifics aside.
- The owner of the project put in place a small (future) mitigation for future forks, but for now IP banlists will probably remain, I'd guess?
- Please try to cache things more aggressively in your build jobs.
[1] https://gmplib.org/list-archives/gmp-devel/2023-June/006162....
> If it builds new images, it'll be at max 10 requests to fetch the latest snapshot in parallel from my end (that's how many parallel jobs Github allows for free), which shouldn't be an issue. I've already switched GMP back to use an outdated release tarball because of how annoyingly flakey their mercurial server is.
Straight from the repo author. Looking at the workflow, not sure how it could really fetch less than it already does.
https://github.com/BtbN/FFmpeg-Builds/commit/d75466340a9a598...
Edit: This seems to be the responsible project/issue: https://github.com/BtbN/FFmpeg-Builds/issues/278
I'm sure it's going to be CI or mirroring or some other automated process.
It's weird to think about github actions crons in the context of this pattern. A casual "I'll mark this to check out later" could result in lots of billable computation.
I often had experiences where people would fork my repos and I would get hopeful that someone might contribute, but it turns out that they just treat forking as bookmarking.
Github running actions for forks wrongly assumes that users use the fork feature for its original purpose.
People already trust GH CI enough to run their code, so why not trust their cache as well?
And remember that users' builds can contain arbitrary custom shell scripts that perform "git clone" or "hg clone" commands. It could be very, very difficult for Github to build an automated system that meddles with the behavior of those scripts to introduce caching, while guaranteeing not to break anything.
Although given how they are using this repo, they'd probably be better off with a shallow clone, or a subversion proxy, or a snapshot tarball...likely just need a set of files..
We're firewalling off all of Microsoft's IP addresses as an emergency response. This is a blunt response, but it is the only response which solves the problem quickly, allowing legitimate site usage to work again."
"UPDATE 2023-06-18: We got a reply from somebody with an impressive title at Github. This person explains that Microsoft and Github have investigated this, and they blame a Github user and the poor GMP infrastructure. It is very interesting that they have done nothing to stop the traffic; we need to keep defending our server by firewalling off more Microsoft IP ranges as I write this 30 hours after Github's response. It is also curious that they blame the victim. (Our infrastructure is pretty resilient with powerful server-class hardware and great connectivity to the Internet.)"
I wonder what % of actions builds are useless runs like this?
Edit: to the shadowbanned guy who hasn’t realised he’s been shadowbanned for years, that’s not correct.
Took me a while to figure out what happened the first time I ran into this.
Most people fork to make a pull request and then never delete the fork after. Why are any scheduled actions required for that?
Cron tasks are the odd exception. You can use them to update dependencies in your repo, or track an external dependency and rebuild your stuff with it, or nightly integration tests with external APIs, etc.
Maybe they could enable all actions but cron actions, but then that’s just as equally confusing. I’d rather GitHub stay out of deciding what’s “expensive” and what’s not. Either enable all actions on forks by default, or disable them by default. They decided to enable them by default and give each user however many minutes of free actions time. I’m sure there is a better upsell opportunity there too.
What?
> Google addressed this issue by having a single enterprise wide cron service.
What???
What does any of that even mean my guy? Neither of these were a cron CI shell script services. What does any of that even mean? Did you just read the word “cron” and get excited?
We had a similar problem. Step one - don't allow submodules from outside of github.com: https://github.com/ClickHouse/ClickHouse/blob/master/utils/c...
If there is a dependency from external git repository, it can be forked to github.com.
But we still have a problem with: - Docker Hub; - Cargo. Sometimes they are offline. Also we should not use them to avoid supply-chain attacks.
The solution probably should be: - copy all docker images as tarballs; - or use ECR. Still not sure what to do with Cargo: https://github.com/ClickHouse/ClickHouse/issues/48575
Shame on everyone who does not appreciate the volunteer work on gmp and ridicules their servers.
The problem started with someone making a poorly built automated build script for another project, and that build script getting forked hundreds of times, triggering a huge load all by the lack of spread in the upstream build script.
This is nothing compared to an actual DDoS attack, which Microsoft was partially blamed for in the previous email. I can see why someone over at MS would be annoyed to find out the supposed DDoS was just 700 requests.
Which is to say... what's MSFT's motivation to worry about this?
And honestly, whoever is maintaining the GMP server isn't appearing particularly competent based on the messages they've written so far. That they'd misidentify or mischaracterize non-attack traffic as abusive seems pretty plausible.
[0] Ok, I'll admit that there is one thing that GitHub should do even if these particular requests are not abusive, just as a best practice when operating a system for running untrusted code that can connect to the Internet. The egress IPs from any single GitHub customer should be sticky, rather than them having a massive pool of IPs shared by all the users. But if that's not how their systems are already architected, making it happen is probably a multi-quarter effort.
However, reading through their response thread [0] I do think they aren't appearing particularly mature. It feels like they are refusing to characterize what is actually going on and digging in and escalating this defensive mindset.
[0] https://gmplib.org/list-archives/gmp-devel/2023-June/thread....
- The fair use update is reactive, and won't help anything
- Not separating microsoft/github from users of the platform (I agree the line is an worthwhile discussion)
- Not following up on the reasonable technical solutions that are being proposed
- Not seeming to ack that while the firewall ban might be the best they can do now, it'll block interested users of their own (GMP) library.
That said, the part in the GitHub reply dismissing GMP's hardware is also immature.
[1] https://gmplib.org/list-archives/gmp-devel/2023-June/006169....
Granlund is a world class expert and apart from gmp has developed (with Montgomery) and implemented several fast arithmetic algorithms in gcc.
This is standard mailing list talk that always includes a bit of posturing. It predates the joyless, grey, boring, judgemental GitHub hypocrites. You are all so mature!
1. They did not provide enough information for anyone at Microsoft reading the message (probably due to being forwarding it to actually investigate the problem. Like, IP addresses, timestamps, the requested URLs. Hell, they even talk of just "identical requests" rather than identifying them as repository clones. Everything about the message is basically custom made to make investigating it as hard as possible.
2. If the goal was to pressure Microsoft into changing something by rousing up a mob with torches and pitchforks, they should have provided enough information for the actual audience (i.e. the mob) to judge the severity of this "attack". Like, the original complaint was about "thousands" of requests. Over what time period? How long was it sustained? (The logs they pasted in the email thread showed 15 requests over a time period of several minutes.)
3. If the goal was to convince Microsoft to take some action due to some reason other than public pressure, than the aggressive and uninformed communications style (copied by their user community) was not going to do them any favors. If you want people to help you, it helps for them to like you or at least have a neutral opinion. This kind of theatrical posturing might make you feel good, but it's going to actively harm your chances of affecting any change.
4. Reacting by blocking traffic from MS IPs to all ports. Given they weren't under attack, there was no reason to block SMTP and potentially block the main communications channel. Or block the website that contains the message they wanted Microsoft to read. But that's what they chose to do. This appears to actually have been like single-digit qps at most. There really was no reason for a layer 3 firewall to be the first solution to try, but even if it were the chose solution you should be able to target it to just hg with less collateral damage.
5. Running an underprovisioned service (they brag about the big real server hardware they have, but then it turns out the repository serving has access to a single core) and then sticking to it. Increasing capacity to soak the problematic traffic is like step one, and would clearly have been an option here. It moves the incident from an emergency to something that could be managed with less stress. And even if IP blocks turned out to be the chosen long-term solution, having a bit of overprovisioning means that things don't break when it turns out you'd missed an IP.
6. Trying to block IPs one by one. It's a server hosting a code repository for God's sake. Literally the first thing anyone should have thought of is that it's a CI pipeline, or something similar, and that if it was MS address space that it'd be GitHub. And the first search result for "github egress ips" will lead to GitHub's official, up to date, list of egress IPs [0], that they could then have blocked in one batch rather than play whack a mole.
So no, despite your snark I think's it's fair to say there was a clear lack of competence in dealing with a potential abuse issue. Given all those signs and the complete lack of details, why would we believe there analysis of this being an attack or abusive traffic were correct? Because, you know, it wasn't.
(Now, is it reasonable to expect everyone to be competent at this? Of course not. But it is reasonable to not just accept at face value the analysis of somebody who clearly does not have the expertise to do it correctly.)
[0] https://docs.github.com/en/authentication/keeping-your-accou...
You should totally move to MS Azure!