CircleCI says hackers stole encryption keys and customers’ source code
techcrunch.com
techcrunch.com
> Zuber said that while customer data was encrypted, the cybercriminals also obtained the encryption keys able to decrypt customer data.
...what? why did this engineer have access to everything? Does CircleCI know what minimum access policies are for?
There should be full transparency on this and all the open source eyeballs should be able to study and scrutinize this. If not, anyone will be able to get away with data theft tomorrow saying, "my server got hacked". What will cause them to not do it except some sense of personal ethics which is rapidly degrading these days?
Private bridges, have owners who give an accounting to no one when they collapse. At best, if you have standing you can sue them and they will defend themselves by giving an accounting for why it’s not their fault or they did their best.
GDPR is a good idea but there must be supervision for it being applied as defined.
Are other companies any better? It seems like now it’s assumed that all of these companies will eventually be hacked, so if you use them you need to have systems in place to mitigate damage. And if you don’t use them them then you need to have mitigation strats anyway. Basically either way you’re screwed.
I'm bringing this up because the circleCI blogpost says that the attacker did memory-dump encryption keys from a running process. See https://circleci.com/blog/jan-4-2023-incident-report/
So even if they were using hashicorp/vault, the attacker could probably still have been able to mem-dump vault's process.
Set up within seconds using a few lines of cloud-init: https://gitlab.com/21analytics/gitlab-runner-cloud-init
Most of the time, it's also cheaper and maintenance is close to zero.
It was easy to take an existing open source terraform module [0], modify it a little for our purposes, and deploy runners that provision job executors in our tightly controlled VPC. It is very simple and everything is open source so you can really understand what it is doing. In our setup, all secrets are issued from our private vault instances. Most of them are short lived or one time use when possible. We are looking at moving our stuff to EKS now as well.
Overall, it was pretty easy to get going but we have the resources to do this. I could see why a small startup would outsource this to someone like Circle CI.
[0] https://registry.terraform.io/modules/npalm/gitlab-runner/aw...
The beauty about CodeBuild is that there is no “lock-in”. All it is fundamentally is a Linux or Windows Docker container with popular language runtimes and a shell script that processes a yaml file or you can supply your own Docker container.
You just put a bunch of bash commands or PowerShell commands in the yaml file and it runs anything.
The Docker container and the shell scripts are all open source and you can quite easily run them locally.
I could see outside of AWS keeping your Docker containers for your specific build environments in a local repository and doing all of your builds inside them using Jenkins using a self hosted CodeBuild like environment.
https://github.com/aws/aws-codebuild-docker-images
https://docs.aws.amazon.com/codebuild/latest/userguide/use-c...
For a “batteries included” approach though, I really like Microsoft Azure DevOps Pipelines.
I’ve even done a couple of integrations between Azure DevOps and AWS when we had clients that are Microsoft shops.
For AWS, if you use CodeCommit (AWS git service), all access is via IAM and granular permissions. If you integrate with Azure DevOps, the AWS credentials do have to be stored in a separate MS hosted credential storage.
CodeBuild also supports at least Github natively.
I’m not shilling for AWS. I have an MS development background (.Net) and only have “DevOps” experience using AWS and Microsoft tooling.
The more Actions matures, the harder it is for them to continue - their market will be legacy projects already on the platform, and perhaps the new projects from companies already using it for older projects (the one reason I can think of to prefer Circle over Actions for a new GH-hosted project).
Defending against internal threat seems like a losing battle that can only be mitigated, slowing down more than preventing an attack.
Circleci's business model is hosting CI runners for you, so of course they need to be able to decrypt the data, and if you have access to dept new infrastructure, you likely have permission to read encryption keys (or deploy new infrastructure that can read said keys and then use those machines to get the keys).
What CircleCI _have_ done is set up enough logging and auditing that they were able to figure out who was compromised, how they were compromised, the time frame and the resources they accessed, which IMO is about as much as you can ask for.
They also point their finger at anti-virus for not detecting the malware. That is a lame excuses. Professional malware developers will check if their product goes undetected by the major anti-virus providers. Anti-virus should not be relied upon to protect against sophisticated attacks
> This machine was compromised on December 16, 2022. The malware was not detected by our antivirus software
That's not blaming that's telling people that their antivirus didn't detect it. If that wasn't there, people would be talking about why they didn't use X antivirus which would probably detect it.
> CircleCI have sufficient logging, however, failed to do fraud analysis to detect irregular access.
This is a significant move of the goalposts. Of course CircleCI messed up, but to go back to the OP's point of "nobody should have that much access to production", well that's just not true.
However, if they need this, you need a detection mechanism. It doesn't have to be advanced; a threshold on the number of production keys accessed per day is sufficient. It should raise the alarm to the manager in question, who can then confirm if this is expected behaviour given the tasks assigned to the employee in question.
Giving access to customer data to your staff by default is a design decision and can be avoided. However, in the end, it is a cost-benefit analysis where you have to decide how much you care about your customer's security. I have worked for enough start-ups/growth companies to know the value put on customers' security is sometimes shockingly low.
This is why Antiviruses are totally useless, they're incredibly easy to bypass.
Not all attacks are targeted and AV can detect abnormal behaviour. Regardless, having the operating system do proper sandboxing is much more valuable than trying to fix it after the fact with AV.
There are lots of tools for key management, but lots of companies don't care about it.
They could run everything in Nitro Enclaves or similar, that require multiple people to deterministically compile and sign new software for, and release secrets into.
I design quorum controlled infrastructure for a living, mostly in fintech where no single human can ever be trusted. You 100% can run infrastructure that, barring a platform 0day, can prevent any single human from having access to the memory of customer workloads and secrets. Customers likewise would encrypt any secrets or code directly to keys that only exist in the enclaves.
CircleCI had negligent security design, but all its competitors are just as bad, to be fair.
Building with security in mind makes you last to market, which is unforgivable in our industry. Getting hacked however is just considered a cost of doing business.
Ominous. What about GitHub? The amount of secrets I trust them with is practically all my life force at this point.
GitHub/NPM have historically failed to support supply chain integrity practices in their public offerings such as hardware anchored code signing, signed code reviews, reproducible builds, multi-party approvals, etc. It is reasonable to expect they are not doing any of that internally either.
Assume any secret you give GitHub will become public knowledge and act accordingly.
The good news is there is never a reason to trust a VCS or CI system with high value secrets. They should never ever need any power beyond running tests, accessing a test environment, or sending notifications.
I can appreciate the want to diversify services so that secrets/env are separate from code, but I think I would honestly trust the behemoth that is Github with both.
That being said, my company still uses Okta, so freebies and mulligans are certainly still tolerated when it comes to data breaches.
That said, I am currently very glad I'm not running anything on circle ci...
A small, 20 person org has maybe 2 people assigned to ops, so monitoring and breach detection is likely worse.
Now, a small org may be a less attractive target and some orgs can have top notch security people, but on average, the trade-off is likely not in favor of hosting your own.
Runners were still self hosted, but if the thing controlling them is just giving 500s all day and you’ve no influence on fixing it, then your jobs aren’t being run and your developers are sitting somewhat idle.
Github Actions has been better, but not perfect… but you can still fully self-host this if you think it’s worth it !
I’ve seen internal hosted Jenkins servers give up the ghost more times than I care to mention.
Personally i’m a big proponent of both cloud services and privately managed services. CI is one of those i think are better kept private due to its sensitivity.
So for instance an employee getting their laptop infected with malware, or through phishing, or through one of the many vulnerabilities discovered regularly on enterprise VPN software?
There is a reason that many organisations are going away with VPNs and "corporate networks" all together - it gives a false sense of security that stuff behind it is protected by the VPN and leads to poor security practices inside like obsolete Jenkins installs.
Disclaimer: I work at a company that sells Zero Trust as a concept and associated software, but I've had this opinion since before joining (I've seen enough of 'there's a VPN and then lots of apps/services/servers with very poor auth practices because it's behind the VPN, why bother?')
Your example of employee laptop getting infected with malware should be remediated by proper device management. Which should include policies like forced updates, prevents software install, uses corporate firewall, up to date antivirus, etc.
Since you mentioned you work for a company that sells zero trust, your response sounds like your solution would replace VPN. Don’t get me wrong, but feels like you are wearing the sales hat now.
Yes, old fashion VPN appliances will be probably seize to exist. Replaced by modern equivalent like tailscale. Layered security is the answer.
From what I hear, Google is moving toward the zero trust direction too.
Google were the ones who started this trend of zero trust
The places I typically work for don’t have the operations capacity necessary. Nor do the small development teams have anyone other than me familiar enough with system administration to set things up in a reasonable amount of time.
A lot of junior developers have little to no Linux server experience. And I don’t really count following a step-by-step tutorial as experience. I have that experience and server problems can easily eat days of time the team can’t afford. :(
More broadly... it shouldn't be that easy to get encryption keys to everyone's secret env variables used for CI jobs.
You cannot shutdown access to resources as developer need them for productivity. Secondly, your competitor is most likely going to take risk of hack and move faster in bringing features while you reduce productivity.
Sort of.
There are ways to verify both the device being used to access a resource, and the account. This is a good use for TPM. One of the features is attestation, which can be used to do this sort of a thing.
For things like crypto keys, this is why HSM (hardware security module) exist. - they make it very difficult to get access to private key (basically, private key doesn't leave the HSM itself, but other cryptographic artifacts based on that key do).
You can also require an authentication loop for each major service, so only services the developer is authenticated to would be vulnerable to credential stealing.
And for critical infrastructure network security would help (e.g. a vpn). With hardware tokens required to initiate the connection, the attacker wouldn’t be able to spin up their own connection.
Finally, access to user secrets shouldn’t be a default level of access for engineers. Instead, that should require a break glass escalation of privileges, reducing the risk of pilfering.
The step-up authentication CircleCI mentions is probably more of a sudo model, with a long-lived baseline level of privilege that can only be elevated for short bursts. This is orthogonal to MFA, but it's at least less of a nuisance with a hardware token than with other options.
> and that site doesn't somehow restrict the session token to only being used on the machine that generated it
One simple practice for security-critical systems is to bind sessions to the client's IP address during authentication. It's not bulletproof given the assumption of malware an attacker can still tunnel through, but neither is locking the session to the device. You could, but the malware is also on the device and can do anything the user can (this is why it's preferable to test user presence outside the device by e.g. tapping a USB key).
1. The engineer was prompted to give the malware (PTX app?) access to the browser key, and agreed to this. Big mistake if so.
2. The malware has an exploit for macOS security.
3. The malware has a way to take over the browser cross-process.
Hardware security keys don't help in this case. What helps is either users not granting malware access to critical secrets, OR, operating system security being enhanced.
A lot of production AWS/GCP keys are likely stolen for those that deploy from CI/CD.
How could a CI company be that negligent. They should be leading this stuff from a best practice point of view.
Security is pretty lax at most companies, and the more employees you have, the worse it is.
Security is lax at most companies but even having worked as an engineer on support rotation, customer anc production access was the last thing I could do and I needed to have exhausted all other options. I don’t feel retaining production access to generate tokens should be a thing.
Malicious files to search for and remove:
/private/tmp/.svx856.log /private/tmp/.ptslog PTX-Player.dmg (SHA256: 8913e38592228adc067d82f66c150d87004ec946e579d4a00c53b61444ff35bf) PTX.app
"Thanks customers for the support" is almost a patronizing thing to say IMO. They should at least offer compensation financially for this and as others have said, his recent update has left more questions unanswered for me.
The way I see it, I'm done as a customer, just need the time to migrate away.
In this case, I imagine that CircleCI's controls involved developer workstations having anti-malware software, logs and audit trails for access to production systems, and some kind of intrusion detection system around those, and some SLAs and policies on how they react to alerts from those systems. It sounds like they had all of those things in place but the malware wasn't detected and their IDS didn't pick up the external access. Unfortunate, but both of those are entirely possible in many environments. SOC2 can't guarantee that your anti-malware systems or IDS are flawless and can't be bypassed by a clever attacker. Otherwise, since they've been able to identify what the attackers gained access to, it sounds like they did have logs and audit trails in place.
SOC2 auditors typically only have a limited understanding of security and technology themselves (the higher end ones might have more resources available to them) and are only able to sample a small amount of data during the observation window.
If you're trusting your sensitive data and credentials to a 3rd party, you should definitely require that they have SOC2 or better, but you still have to own your own security and consider how you need to protect yourself if/when they are compromised.
The method of attack sounds like CircleCI's production cloud (probably AWS) was impacted - "the targeted employee had privileges to generate production access tokens as part of the employee’s regular duties, the unauthorized third party was able to access and exfiltrate data from a subset of databases and stores, including customer environment variables, tokens, and keys."
But I am surprised that their SOC2 auditors didn't raise exceptions about their lack of controls. Sounds like a pretty immature program, they only talk about 2FA, MDM and SSO which is basic stuff. Where is the SIEM? Or CSPM? Or any alerting!? Yes there are SOC2 automation platforms out there that rubber stamp stuff, but at CircleCI's scale I'd expect more scrutiny.
The extent to which CircleCI has gone to eliminate all threats is... scary. They've gotten GitHub to invalidate any GitHub access tokens used by a CircleCI customer. They've gotten AWS to e-mail AWS customers if one of their access keys was stored in CircleCI. It's a complete and total compromise of literally every customer secret in CircleCI. I expect this will be the biggest hack of 2023... and it's still January.
So, yeah, I'm pretty sure customer source code is up for grabs.
> Review GitHub audit log files for unexpected commands such as ... repo.download_zip
> ...
> Review GitHub audit log files for unexpected commands such as ... repo.download_zip
I don't know what you'd get from GitHub other than source code but the CircleCI blog post explicitly describes the attacker downloading entire repos as a .zip
But getting the source code is pretty bad by itself. Not from an intellectual property standpoint, but because I've never seen a company whose developers didn't commit live credentials into their source code.
Said company ought to immediately rotate the credentials and force rewrite the repo history to nuke the commit for good measure.
I don’t do that, but I’m pretty pathological about stuff like that.
I learned it from the company I used to work for, who were paranoid to the point of lunacy, about Chinese hackers (they were breached once, and took the lesson to extremes).
I don’t think they are an outlier. I’ll bet lots of companies are just as tinfoil.
It’s a downright unbearable development environment, though.
I’ve heard banks can be even worse.
> Updated headline to better reflect the customer data that was taken.
Probably should be updated on HN too.
I haven't seen much discussion on how this specific attacker entrypoint can be mitigated. So I'm going to make a naive attempt in this comment.
How about storing the client's IP address in the session cookie. Then whenever the server recieves the cookie, it compares the client's IP address against the one stored in the session cookie. The server denies the login if there's a mismatch. The cookie would of-course have to be signed(hmac etc) so that it is tamper proof.
One problem with this is that client IP addresses are easily spoofed[2].
So, instead of storing the client's IP address; how about we instead store the clients' SSL fingerprints[3][4]. I haven't looked much into the literature, but I think those fingerprints are hard to spoof.
1. https://circleci.com/blog/jan-4-2023-incident-report/
2. https://adam-p.ca/blog/2022/03/x-forwarded-for/
That doesn't work in environments with multiple NAT origin IPs in place, or when they're using crap like Netskope/some other "security"/"privacy"/"VPN" software, as IPs tend to randomly change with these. It would generate way too many false-positive reports.
> One problem with this is that client IP addresses are easily spoofed[2].
Only if the backend servers are badly set up. For me, I always run haproxy as the frontend and forcibly delete incoming headers (X-Forwarded-*, Forwarded), and as an added precaution the backend software is configured to only trust the haproxy origin IPs - so even in the case an attacker manages somehow to directly access the backend servers directly, they cannot get a spoofed IP past the system.
> So, instead of storing the client's IP address; how about we instead store the clients' SSL fingerprints
That requires client-side SSL authentication, which is theoretically supported by all major browsers, but very rarely used and the UI support is... clunky at best.
I do not think it requires client side SSL. See: https://engineering.salesforce.com/tls-fingerprinting-with-j...
What is been fingerprinted is the TLS negotiation between client and server.
Also, maybe that localised build system running on an old server for each team seem to be not a bad idea to reduce the blast radius when eventually a hack happens. These providers are supposed to be the gatekeepers and experts who one leaves the tedious and critical work to. If they are just being a leaky cauldron, maybe not bad to cook in my old pot at home.