Remote code execution vulnerability in Google they are not willing to fix
giraffesecurity.dev
giraffesecurity.dev
It's really unfortunate that even newer package management ecosystems like Rust's haven't learned that lesson. Yes, it increases initial setup friction, but dependency confusion and name squatting issues more or less just go away when you do that. Sure, you might run into issues where people register domains for common misspellings of popular domain names, but that's already a problem with website resolution, and there are ways to mitigate that.
I thought this was the expected behavior of --extra-index-url
To force pip to use a user-specified registry, you would use --index-url instead of --extra-index-url
If PyPI used Java-like scheme with domain ownership verification, Google would presumably use com.google.* namespace for their private stuff and it wouldn’t be possible to introduce malicious package this way.
So the described issue arguably is a consequence of that.
That there is no other option for private repositories or namespacing is still a big shame.
I don't know the policies on what would happen if someone else registered the domain and contested the ownership though.
Node.js is particularly notorious for this. It’s a big challenge for having non-flat dependency management and there’s no great solution for this problem.
All so someone can save few letters with typing dep name that one time they will be pulling it in into their project
They aren’t the only one, but they’re… prolific in this area. A list of squatted names is maintained here: https://github.com/dtolnay/squatternaut
Requiring that third party dependencies be signed should have been the default in Maven itself.
A package called com.example.foo would be package foo in the namespace "example", which is under the namespace "com", which is how a lot of languages do namespace nesting.
DNS can be said to be "reversed" from that point of view.
Chesterton's fence.
That doesn't really solve the fundamental problem, though. Domains can and do change hands by various means and for various reasons having little to do with Java packaging, meaning that control of those packages can be transferred as well. It's just shifting the problem to a different place is all.
Obviously not true, in fact none of the companies I worked in that was the case. Now, if we're talking about somebody like network admin that may be more juicy, but in a large company rarely a single person has the "highest possible level".
> If you never worked at a company that took it seriously it is hard to imagine that there are people who do take it seriously
This sounds quite unsufferably smug, and unwarrantedly so. Yes, we get it, in Google they have access levels. Newsflash: this concept wasn't invented by Google. A lot of companies use it. Running code on engineers workstation does not give you keys to the whole castle. But it does get you into a first perimeter, after which you can do some recon, not available outside the big firewall, launch other things against corporate sites not available to mere mortals, look for "experimental" servers which are poorly secured because they're not production yet and are behind corporate firewall (that happened in pretty much ever company I've worked with), get their hands on some juicy browser cookies containing authorizations to some internal services, steal some code finally... There are a lot of interesting things a person can do with engineer's access and some of them may land a stepping stone to a next step access maybe. If you don't understand why such a thing may be important even without giving somebody the keys to the castle - maybe you don't know enough about how secure systems are built to smugly dismiss other people.
Selection bias. In some companies access to an engineers single workstation allows access to potentially billions of dollars worth of intellectual property. Imagine China stealing the design for the latest aircraft engine from GE, or accessing the PC of the most senior accounts payable person in a company and getting access to actual money. Just because YOU weren't the most valuable target in a company doesn't mean no one in your role is. My account at my company is constantly under attack because I'm the VP of IT in healthcare. The reality is "my" account doesn't even have any powers, it's just email. My accounts with sensitive access are separate.
> People [..] freak out about this because [..] in 99.9% of companies, running code on an engineer's workstation would immediately be the highest possible level of breach.
So it's not selection bias, it's a counterargument. The poster also said engineer not "VP-level".
So, your comment is not really relevant.
I once offered a bet to the large security team at a well-known decacorn tech company I worked at: I offered to make a personal, reasonable-sized cash bet with any member of the security team that I would win if I could deploy malicious, unreviewed code to any service or machine of their choice without it being prevented or proactively noticed by them.
The members of the security team all declined my bet. We're talking about a team of probably at least a dozen people, many of who had been working at the company far longer than I and who had been shaping and reviewing the company's security design for years.
They knew perfectly well that I would be able to win the bet. Not because their security was unusually bad, but because it was bad in the common, usual ways. Securing the supply chain is hard, and real security is almost impossibly expensive to add to a system late in the game if you didn't design it in from the beginning.
It's fine to not be confident, but when professional security teams at large companies are afraid to express confidence that their systems are non-trivial for a random engineer to hack in their free time, that seems at odds with the claim that it's "obvious" that permission escalation is hard
A willingness to take pride in your work and to not take it too seriously when smart, well-intentioned people make mistakes (e.g. blameless postmortems) is part of the culture difference that led to Google's engineering becoming so exceptional and innovative vs the more corporate, don't-rock-the-boat, fear-driven culture that the traditional businesses had at the time.
I'm assuming you were at google in late 90s/early 2000s?
I've long thought that one should have the attitude (and act to make it so) that one should be willing to bet their job on the quality of their work, but not necessarily actually do so.
And betting anyone (co-worker or not) that they can't compromise the systems (especially, but not limited to production systems) you're tasked with keeping from compromise is a bad bet -- even if you win.
I'd class that sort of behavior as having serious potential to be a "Career Limiting Move" (CLM).
There is no big firewall to bypass here. That's the whole point of zero trust.
Also, about “code reviews”: here’s a story from last year about how a massive refactoring of some webkit code in 2016 resurrected a massive exploit that was actually fixed in 2013, but went unnoticed. It was only discovered and patched in 2022, as it was being exploited in the wild by, among others, the NSO group.
https://www.theregister.com/2022/06/21/apple-safari-zombie-e...
Even if you only take over one developer's system, it's a great starting point for pivoting into the network and starting a more sophisticated attack. I'm sure an advanced threat actor would know how to take advantage of the opportunity against Google.
We both know that Google does better than most at endpoint security. In some cases it’s possible to argue that they are the best. What we definitely don’t need is you to be superior about it: it’s part of the reason why (ex-)Googlers have a poor reputation.
In this case, having technical measures that avoid sensitive things ending up on developer machines is an excellent way to help improve your security posture. That said, it definitely doesn’t mean you shouldn’t be unconcerned about code execution on developer machines. There’s a reason that internal red team exercises distinguish between external access and already having a foothold on a machine.
In this case isn’t obviously not a specific failure in Google’s security policy that package managers don’t do namespacing, but if I was on the team and I received this report I would at least think about whether there is something I would want to do here to improve the situation, similar to existing efforts to prevent attacks like paste bracketing or trivial keylogging.
Seems weird to group those two as the same thing
tl;dr - if I were still gLinux security, I might not be freaking out about this, but it would definitely fall into the set of stuff I'd be making space for in next quarter's OKRs.
Haha I hope this to be true :sigh: The reality is that all those security hardening measures already surpass the level where it significantly undermines the overall productivity... Engineers cannot even have a test run on production data without an explicit review from colleagues.
Tools like code search and the source checkout process also both check for accessing unusually large portions of the codebase, making it only possible to exfiltrate small portions of the codebase at once.
https://github.com/google/santa
This is a product developed by Google that has at least been utilized internally to some extent. It's not perfect, but my previous company used it and it does prevent unexpected unknown code from running in the background.
What it does not do is prevent someone from intentionally downloading and executing a library unless the upvoter actually comes to some demand that you do so. I found that it quickly became a bit of a "alert fatigue" where you approve things your coworkers send you so they can get back to work without properly vetting.
In a well designed zero trust network this makes little difference. The traditional posix security model is bogus from the start anyway. A lot more useful stuff you can exfiltrate as a regular user usually.
So just like pretty much any package from public repo out there ?
(I work at Google but have no special insight. My opinions are my own.)
(Or, is this comment https://news.ycombinator.com/item?id=35585453 complete and accurate?)
Surely that is not the only viable defense?
“A project published on the Package Index meeting [...] the following is considered invalid and will be removed from the Index: [...] project is name squatting (package has no functionality or is empty);”
although the enforcement of that rule is lax in general, not only in this instance[2].
[1] https://peps.python.org/pep-0541/#invalid-projects
[2] e.g. https://pypi.org/project/requests3/ off the top of my head
* There is no namespaces, so all internal packages must be individually protected * It's easy to misconfigure the various python package tools to open you up to dependency confusion attacks * The only effective protection mechanism that you can implement across a large enterprise fills the index with spam and is forbidden
It would really improve things if they could introduce namespaces and let legal entities own those.
That’s not a perfect solution, but so far that DNS is the namespace with mainstream acceptance and builtin lawyers that a single entity cannot e.g. just singlehandedly sanction, sue or simply “reserve the right to refuse” people out of—in practice, for the most part.
PyPI’s (and CPAN’s, CTAN’s, Hackage’s, NPM’s) centralized index was originally a (deliberately) crude solution to the discoverability problem, at least in part. These days, we have adopted a different bad solution—putting everything on a Microsoft-owned hosting service with a crap search function. That is also quite bad, but maybe it’s time we recognize it happened anyway and stop making concessions to the old solutions in our package naming schemes.
Certainly it isn't "squatting" in the typical sense of "prevent someone else from using a valuable name".
Package registries need to address this, as even the newest and "modern" ones keep repeating the same mistakes (yes, I'm looking at you cargo/crates).
Just add a damn namespace, where packages have to be under in order to be publicly available. Would solve the problem yesterday, but instead new registries with global names keep appearing like it's not a problem.
In the case of heavy moderated ones like Debian et al, it makes sense with a global namespace, but for the ones anyone can upload a package? Require namespaces already...
https://learn.microsoft.com/en-us/nuget/nuget-org/id-prefix-...
NPM has scopes:
https://docs.npmjs.com/about-scopes
So there are ways to make package repositories prevent squatting on package names.
It doesn't look like the fixing effort is progressing very quickly: https://github.com/pypa/pip/issues/8606
To their credit, at least they didn't close it "works as intended" which I imagine a lot of projects would.
It is working more or less exactly as it should, the employees in question are just downloading untrusted packages inadvertently. This is something that could happen anywhere, even without private package repos.
I think my response would be two words and the second one would be off.
It's actually deliberately criminal as I read it. "Hey, I trojaned some code and got it downloaded onto your company's systems! Please pay me a bug bounty!" is 100% isomorphic to extortion.
There is no credible threat here. In addition to the above points, if attacks were made by exploiting these facts, the author, having raised the issue in the first place, would become a person of interest in any investigation.
Notice that you are also taking the author at his word when you say he literally compromised live systems. To turn this into a case of extortion, you would have to go beyond that and invent a number of things that have not been said - and some highly implausible things at that, given the very public way in which this supposed extortion is being conducted.
Technically true. If this happened as described, it's criminal behavior. If he's lying, it's maybe not.
FWIW: explicit threats are not and never have been a requirement for prosecuting extortion crimes. I'm not sure where you are getting that.
And for your information, 'explicit' is not a synonym of 'credible'.
The issue here is that the submit actually attacked live systems, instead of just reporting on the possibility of malicious library code.
...which is something everyone already knows about, and thus why he couldn't get paid. You don't get paid for actually hacking systems either!
It is a google vulnerability for using the tool in ways that are known to be broken. Dependency confusion attacks are well known and have known mitigations. When depending on private packages one must not rely only on extra-index-url, instead point to a full url or use a completely internally hosted index-url.
I mean, this guy just registered a package without mentioning it to anyone, then suddenly it started executing inside a google users machine. No social engineering involved. Note that pip install is not just downloading, it also can run arbitrary code during installation phase due to setup.py.
Sure, Google likely have another 3 layers of defense to get to the truly interesting sauce, at least he got through the front door.
The two guys responding to the email are basically "doing their job". "Oh, it's not a bug in the software package, no cookie for you." Yeah, it's a much more severe incident, you stupid son of a bitch. Any CISO sees this and throws themselves out of the window.
You obviously had a good point here but including things like "you stupid son of a bitch" unfortunately flips a higher-order bit. Perhaps you don't owe inadequate vulnerability handlers better, but you owe this community better if you're participating in it.
If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
At Google’s scale you need to assume that even some employees will be bad faith actors (e.g. agents of some government, with a goal of surreptitiously adding back doors) and you need far more sophisticated security controls (e.g. multi-party controls, immutable audit trails) than assuming engineers or their systems will never be compromised. The latter is going to be true for some employee nearly 100% of the time even if you don’t have bad actors.
The existence of these controls and general set of security assumptions and architecture are what make this not a big deal, not a lack of care.
Internal package registry that knows all internal package names and makes pip reject colliding names from other sources would be another possibility.
Explicitly verifying the hash of the internal package based on the registry (again) and refusing to install packages that don't match the hash would be another option.
I'm sure if a person smarter than me (Google probably has thousands) spends a day thinking about it, they could think of a dozen better ways.
come the fuck on
I don't know where some HN folks get their logic from.
The possible ways to respond to that process are open ended and infinite. You can do anything you want about it. Doing nothing about it is approximately the least defensible.
Everything you do has a cost. Increased friction which makes people work around it. Lower productivity. Time spent implementing it that cannot be used to implement something more useful.
Don’t just do something. Do something that improves the situation - and google does. Their statement indicates that they consider the developers machine as fundamentally not trusted - and I’d consider that a correct assumption. Some of the thousands of machines will be compromised at any given time. Whether it’s via this exploit, or another or by bribing the engineer doesn’t matter. What matter is that they attempt to contain the issue at that boundary.
I did not say or imply or suggest to "do something, anything" without caring if it's sensible or effective.
"anything" simply means there is no limit to the possible suitable things.
There is also no limit to the possible unsuitable things, but so what?
It is beyond stupid to take that starting point and conclude that anyone suggested "Maybe Gooogle should issue Tarot decks to all employees to determine if they should press enter at the end of every shell command." just because, after all, that is something and included in "anything".
I can't know which of the infinite possible detailed measures make sense within Google's environment. But I don't have to to still know that they exist. The details will depend on internal details only they know.
If I say "wrap the pip command in an internal wrapper that performs various checks" surely there is some reason that is not practical or not effective enough, exactly as stated. Or maybe that would exactly clear it all up. But if not, ok so something else then. Have some imagination. But that does not remotely imply random nonsense.
You can protect production and CI systems by restricting their internet access, but who can reasonably do work with such a restriction.
First, it’s not required to actually run the package. Installing it is sufficient. All package managers that I worked with so far can run code at install time.
Second, the issue here seems to be a misconfiguration that makes pip look up a private package that should be retrieved from a private repository on the public repository. The attacker then just registers a malicious package on the public repo with the same name. Preventing this attack requires that python is correctly configured on each an every developer machine - something that I’d never rely on as cornerstone of my security.
Third: This is one example of smuggling a malicious package on a developers machine. Another vector is that a good package turns into a malicious package with an update. That’s even harder to defend against - pulling in the update with the programming languages standard tooling may run malicious code. You can certainly first download the package, unpack and inspect it and then pull the update - but would you rely on thousands of developers diligently doing that?
Last: this class of error affects almost all programming language package managers out there.
So it’s better to assume that this will happen, take a local compromise of a developers machine as a matter of when, and mitigate what the attacker can do with the capabilities they gain from this compromise.
True enough. The problem is that a) it is a common misconfiguration and b) it appears to be also affecting Google, which is a big juicy target for any computer criminal. We're not talking about solving a theoretical problem in 100% of theoretical cases. We're talking about having a very practical vulnerability - which can be practically fixed. Yes, that doesn't fix all other theoretically possible vulnerabilities - so what? That's like arguing that since halting problem is unsolvable having debuggers and static analyzers is useless - we can't solve 100% of the problem, so why even bother to solve even 1%?
> Last: this class of error affects almost all programming language package managers out there.
Again, you're replacing a specific issue with a "class". Yes, you can't fix all the problems in the whole class. But you can very well fix this particular one, in many ways.
Sandboxing the development environment could be done, but would only help against this attack if the sandbox cannot connect to the public internet- which again would be painful.
Googles strategy of accepting that this kind of breach will happen and rather focus on mitigation of the resulting damage seems like the better way.
We fixed* this right away, because even though it's true that this "vulnerability" exists with basically every npm package, the difference is that anyone can immediately pull this off once they find an unpublished package in use - they don't have to take over an existing package or get a package they own to be used.
It's the ease of the executing exploit that makes this one more dangerous. Some bored kid could have just wiped my hard-drive or worse, maybe within a few minutes if I'm working.
* An easy fix on npm is to create an org that you use for all internal packages. No one else can publish to that org.
From the article [1]:
Rationale: Code execution on a Googler machine doesn't directly lead to code execution in a production environment. Googlers can download and run arbitrary code on their machines - we have some mitigations against that, but in general this is not a vulnerability; we are aware of and accepting that risk.
[1]: https://giraffesecurity.dev/posts/google-remote-code-executi...
All I am saying about the issues raised in the article is that some organizations start from the assumption that engineers will run arbitrary codes on their devices and the rest of their security story follows from that. It is not necessarily irrational. If you think that is crazy, it might be because you don’t understand the base assumptions.
https://pip.pypa.io/en/stable/cli/pip_hash/
Apparently the format looks like this for a requirements.txt file:
# sha256: L9XU_-gfdi3So-WEctaQoNu6N2Z3ZQYAOu4-16qor-8
drf-compound-fields==0.2.0
FooProject == 1.2 --hash=sha256:2cf24dba5fb0a30e26e83b2ac5b9e29e1b161e5c1fa7425e73043362938b9824
Follow the link "Hash-checking Mode" from the page you linked.Seems like you should? It doesn't even have to be malicious, but you're much more likely to get a response (even if it's just a bug fix) by opening a window on someones computer that says "haha, you're hacked! send and email to security_team1234@google.com and let them know what happened."
caveat: not in infosec – maybe there is precedent to not do this kind of thing if you're in the business of bug bounty hunting.
Part of the reason we have a crisis in computer security is because the good guys have to be extremely careful about the systems they poke. They can only poke companies with responsible disclosure policies in specific ways. It shouldn't be a crime to find and report vulnerabilities in good faith, but that's how it is. I almost got myself in big trouble for doing so on one occasion.
Meanwhile the actual bad guys are getting away with draining bank accounts and dumping databases with millions of peoples' personal information.
A package whose README says:
"This is an automailer to send the CEO of Google a friendly Hello, and politely request changes to the Bug Bounty program."
Which does exactly that, and nothing nefarious beyond that, would probably be okay. It's doing exactly what's advertised.
You want to avoid anything which uses words like "hack" or "compromise." Indeed, you can go out-of-the-way to point out it is explicitly not a "hack" or "compromise" under current Google policies.
"Something malicious" would be very different than sending a proof-of-concept email. "Something malicious" might be, for example, snarfing up data, or having one engineer commit malicious code and having another one approve it.
Indeed, the email could walk through malicious use-cases like these, which either leak customer data or damage Google infrastructure.
Google must be really confident in their ability to contain threats like this one. In other orgs, this would be a "hair on fire" sev 1 security incident.
Hackers “discover” a password, but infrastructure is watching for that canary to sing, and they know exactly what machine or file was compromised.
Hackers “discover” an internal service with a known vulnerability, but attempting to use the apparent vulnerability (honeypot) triggers security.
Dependency Confusion: How I Hacked Into Apple, Microsoft and Dozens of Other Companies The Story of a Novel Supply Chain Attack https://medium.com/@alex.birsan/dependency-confusion-4a5d60f...
A company the size of Google should have been prepared.
left-pad's author could have done anything with all the dependencies on his package.
Downloading from local is more natural, imho.
In Java builds we usually had:
build -> org local repo -> maven central
So the local repo (be it Artifactory or Apache Archiva) works as a proxy. It dowloads the artifacts form internet if the artifact is not present locally. The build does not go directly to the maven central.In such environments, it is crucial to maintain a high level of isolation between development machines, personal computers, test farms, production servers, and source code repositories.
Also there are various types of development machines and environments to cater to different needs. Some are highly restricted, only allowing developers to interact with and modify the production source code. These environments provide minimal functionality beyond code editing and submitting changes to test farm. On the other hand, there are more flexible environments, akin to "scratch pads," where developers have greater freedom to experiment and explore.
I think this is something GoLang does a great job at. To use a package, you have to specify the exact URL of the repo. This mitigates the risk of dependency confusion since an attacker would need control over the domain to upload a conflicting package.
For distribution repositories (apt and so on), yes there's some vetting. To start with, only a limited set of people (the distribution's developers) are allowed to upload packages to a distribution repository, and even then, there's often a second layer of vetting for new package names. For instance, on Debian (https://wiki.debian.org/Teams/FTPMaster):
"When a package is uploaded to the unstable or experimental suite, it falls into one of three categories. If it is a new version of an existing package and adds no new binary packages, it is moved into the package pool automatically. If one or more of the binary packages or the source package itself is not currently in the archive or if a package is moved between the components (main, contrib, non-free), it is NEW and must be examined by an FTP Team member (see NewQueue). [...]"
That is, when a new binary package is added, either because it comes from a new source package, or because a source package was modified to add a new binary package, the FTP masters have to manually approve it, before it becomes available to be installed by apt.
This is different from language repositories (pip and so on), in which anyone can register a developer account, and there's no manual vetting of new package names.
default apt repos are vetted by Debian/ubuntu folks.
so its understandable that they closed it as non-vulnerability.
That's what I thought. Until I thought a bit more.
Firstly he didn't fool the trusted human into doing anything they weren't already doing anyway. I don't see anyone being tricked.
Secondly, a lot of exploits depend on someone doing, unprompted, something they really didn't ought to do. I mean, if you rule out as an exploit anthing that simply depends on people not taking active measures against an attack they didn't know about, there's not a lot left.
I don't think Goo should have paid out, though; there's no vulnerability shown in any Google software, whether a product or an internal tool. It looks to me like a pip vuln that $TRUSTED_HUMAN could and should have evaded.
Do they not realize that most big tech companies have moved on to single feeds that are governed by their own security/inventory teams? Using public registry is an anti pattern now and has been for awhile, well before “dependency confusion”.
Not all package managers have implemented a stopgap to the problem either. I’m a bit disappointed to see this article though. The world runs on trust and we all trust that people won’t abuse known vectors for their own gain.
He should have it install a reverse ssh tunnel then pass along keylogging and a screenshot every 2 seconds, he'll likely find someway to pivot for a 'vulnerability'.
What a joke google is.
Way to go HN, if people know this little about security than this place has become a complete cargo cult. Enjoy your false idols.
I've banned your account just now, however, since you don't seem to be using HN in the intended spirit.
If you don't want to be banned, you're welcome to email hn@ycombinator.com and give us reason to believe that you'll follow the rules in the future. They're here: https://news.ycombinator.com/newsguidelines.html.
But of course Google has large stockpile of IPs already so they are not really impacted the same way others might be.
2. It relies on tricking a developer into downloading a substituted package, which is indeed social engineering.
3. If google were suceptible to malicious code execution machinations of singular employees in China, Huawei would be google, not google.
I wouldn't call this social engineering. The attacker isn't actively trying to trick anyone of anything. They're just exploiting the fact that the Python package management tools make it really easy for a user to accidentally -- without any prompting or interference from the attacker -- pull packages from pypi.org rather than their internal private repository.