Apt Encounters of the Third Kind
igor-blue.github.io
igor-blue.github.io
Reading between the lines, I sense the client didn't 100% trust Mr. Bogdanov in the beginning, and certainly knew there was exfiltration of some kind. Perhaps they had done a quick check of the same stats they guided the author toward. "Check for extra bits" seems like a great place to start if you don't know exactly what you're looking for.
Their front-end architecture seemed quite locked down and security-conscious: just a kernel + Go binary running as init, plain ol' NFS for config files, firewalls everywhere, bastion hosts for internal networks, etc. So already the client must have suspected the attack was of significant sophistication. Who was better equipped to do this than their brilliant annual security consultant?
Which is completely understandable to me, as this hack is already of such unbelievable sophistication that resembles a Neil Stephenson plot. Since the author did not actually commit the crime, and in fact is a brilliant security researcher, everything worked out.
This is hardly reducing the attack surface compared to a good distro with the usual userspace.
It's been decades since attackers relied on a shell, or unix tools in general, or on being to write to disk and so on: it's risky and ineffective.
Many attack tools run arbitrary code inside the same process that has been breached and extract data from its memory.
They don't try to snoop around or write to disk and so on. Rather, move to another host.
The only good mitigation is to split your own application in multiple processes based on the type of risk and sandbox each of them accordingly.
Run `tcpdump -n 'tcp and port 80'` on your frontend host and you'll still see PHP exploit attempts from 15 years ago. Not every ghost who knocks is an APT. A singleton Go binary running on a Linux kernel with no local storage is objectively a smaller attack surface than a service running in a container with /bin/sh, running on a vhost with a full OS, running on a physical host with thousands of sleeping VMs—the state of many, many websites and APIs today.
No, you have to understand what is really part of the attack surface and what the attacker wants.
For example, on a properly built system with a single application running with its own user the attacker might have no practical benefit at all in doing a privilege escalation to root.
> running on a physical host with thousands of sleeping VMs
This is a strawman. A shared hypervisor opens another attack surface and was not part of the discussion.
EDIT: Anyway, I'd like to thank Mr. Bogdanov and his client for sharing this story—it's just fascinating.
Not to worry. When the ghosts knock, you just have to remind them that their attack is out of scope /s
Unless you are running unnecessary daemons exposed on the Internet, 99% of the attack surface is from your application and the kernel itself.
Both parts that you can't remove.
If you suspected your security consultant, what would be the point of slipping them tiny hints about what you've found? If they're the source of the intrusion, they already know. If they're not the source of the intrusion, why fear them when you've already been compromised? Also, if you suspected the consultant, why hire them to do the security review?
I suspect the real reason is probably simpler: they have strong personal or financial incentives to "not have known" about the intrusion before the researcher discovered it.
Specifically I don't think the owner thought it was likely, just a concern he couldn't shake. Probably he relaxed as soon as the consultant didn't make excuses, and tackled the job—extracting the binary from an unlinked inode is definitely not showing reluctance. Pure speculation, of course.
Why is it necessary to point out the foreign origin? Doesn't that just encourage our innate xenophobia?
To me your interpretation reads like: something is wrong, must be (substitute with your enemy du jour)
For instance, sometimes I say something to a friend and they misunderstand what I intend. The feedback on the misunderstanding permits me to recalibrate my communication and it helps them receive the right information.
I am not claiming that it was "the Chinese". I'm claiming that saying "Chinese APT" reads to me like this.
However, I'd propose a new convention that any unattributed attacks and example threat scenarios of nation states should use Canada as the default threat actor, because nobody would believe it or be offended.
Meanwhile, can you prove that this "innate xenophobia" is present in every human to an extent that it's actually relevant, and that this particular instance of suggesting that the malware is Chinese in origin meaningfully exacerbates it?
Moreover, China is a geopolitical rival to the United States, India, and other countries that constitute a majority of HN readers. Information like this is interesting from that viewpoint.
To prove my point s/he had no problem with the top level comment 6 hrs ago “mossad gonna mossad”
But when you have what appear to be state level actors using 0 day exploits... you will not stop them.
A packet capture of the config files would show something was up to anyone suspicious, but knowing what to do about it is a completely different story.
Of course, I have no non-circumstantial evidence and this could all be a coincidence, which is why my comment is prefixed with "conspiracy theory".
1: "However, he asked me to first look at their cluster of reverse gateways / load balancers"
2: Would have likely been less likely to find the issue with active analysis given the self destruct feature
3: "Specifically he wanted to know if I could develop a methodology for testing if an attacker has gained access to the gateways and is trying to access PII"
4: "I couldn't SSH into the host (no SSH), so I figured we will have to add some kind of instrumentation to the GO app. Klaus still insisted I start by looking at the traffic before (red) and after the GW (green)"
Perhaps they had noticed the programs restarting and when trying to debug triggered it.
The other ones could be explained by him being afraid of leaking PII, and most PII being on that system.
Not wanting to instrument the Go app could be an operational concern.
"Your <reverse gateway> devices are compromised and leak PII." Nothing more.
>I think he had some suspicions, but he is denying that vehemently ;)
https://twitter.com/IgorBog61650384/status/13753134251323146...
How the attackers were able to gain access again after the developer used a VM in Windows? My guesses:
- The developer machine was compromised in a deeper level (rootkit?)
- The developer installs a particular application in each Linux box
- There is a bug in an upstream distro
Unlikely that would not have taken 3 month.
> The developer installs a particular application in each Linux box
Possible, but also unlikely, as long as the vm wasn't used for other things this also wouldn't have taken 3 month.
> The developer installs a particular application in each Linux box
There probably is, but it probably has nothing to do with this exploit. For the same reasons as mentioned above.
My guess is that it was a targeted attack against that developer and there is a good chance the first attack and the second attack used different attack vectors hence the 3 month gap.
Steganography begs to differ.
How much free entropy do you have on your network traffic?
EDIT: Corrected. Thanks cuu508.
(The CI was not compromised but a dev laptop which was used to manually build+deploy the kernel, without any CI involved).
Through generally I agree with you.
If you roll your own build environment then automate the build process for it and recreate it from scratch fairly often. Reinstall the OS from a trusted image, only install the build tools, generate new ssh keys that only belong to the build environment each time, and if the build is automated enough just delete the ssh keys after it's running. Rebuild it again if you need access for some reason. Don't run anything but the builds on the build machines to reduce the attack surface, and make it as self contained as possible, e.g. pull from git, build, sign, upload to a repository. The repository should only have write access from the build server. Verify signatures before installing/running binaries.
And I guess, for those super-critical builds, don't rely on anything but the distro repos or upstream downloads for tooling?
Because if you deploy your own build tools from your own infra, you are at risk to taint the chain of trust with binaries from your own tainted infra again. I'm aware of the trusting trust issue, but compromising the signed gcc copy in debians repositories would be much harder than some copy of a proprietary compiler in my own (possibly compromised) binary repository.
You can build more tooling by building it in the trusted build environment using trusted tools. Not everything has to be a distro package, but the provenance of each binary needs to be verifiable. That can include building your own custom tools from a particular commit hash that you trust.
I feel reasonably competent at defending my code from attackers. The stuff that runs underneath it, no way.
That can not be the right lesson, because there's no inherent reason "personal machine" is any less safe than "building cluster" or whatever you have around. Yes, on practice it often is less secure to a degree, so it's not a useless rule, but it's not a solution either.
If it's solved some way, it's by reproducible builds and automatic binary verification. People are doing a lot of work on the first, but I think we'll need both.
Sure there is! I browse internet a lot on my dev machines, and this exposes me to bugs in browsers and document viewers. And if I do get compromised, my desktop is so complex and runs so many services the compromise is unlikely to be detected. So all attacker needs is one zero day, once.
Compare this to a CI with infra-as-a-code, like Github Actions. If the build process gets compromised, it only matters until the next re-build. Even if you get a supply chain attack once (for example), if this is discovered all your footholds disappear! And even if you got the developers' keys, it is not easy to persist -- you have to make commits and those can be noticed and undone.
(Of course if your "building cluster" is a bunch of traditional machines which are never reformatted and which many developers have root access to, then they are not that much more secure. But you don't have to do it that way.)
Securing the machines themselves is a process of adding up always decreasing marginal gains until you say "enough", but the asymptote is never towards a fully secure cluster. That ceiling on how secure you can get is clearly suboptimal.
Besides, the ops people's personal machines have a bunch of high access permissions that can permanently destroy any security you can invent. That isn't any less true if your ops people work for Microsoft instead of you.
Your build actions happen all in the docker images/ephemeral VMs. You use images directly distributed by the corresponding project, for example you may start directly from Canonical's Ubuntu image. The "runners" are provided by Github, and managed by Microsoft's security team as well. The only thing that you actually control is a 50-line YAML file in your git repo, and people will look at it any time they want to add a new feature.
Yes, the if someone hacks Microsoft's ops people, they can totally mess up my day. But would they? Every usage of zero-day carries some risk, so if attackers do get access to those systems, they'll much likely to go for some sort of high-value, easy-money target like cryptocurrency exchanges. Plus, I am pretty sure that Microsoft actually has solid security practices, like automatic deployments, 2FA everywhere, logging, auditing, etc... They are not going to have a file on CI/CD machine that is different from one in Git, like OP's system did!
The APTs do not have magical powers, they buy from the same exploit market everyone has.
Let’s say my organization (which is not very well known) has an exploitable bug. What are the chances that someone will discover it? Pretty close to none, the hole can be there for many years waiting for APT to come and exploit it.
Now imagine Github runner or default Ubuntu image has an exploitable bug. What are the chances it will last long? Not very high. In a few months, someone will discover and either report or exploit it. Then it will be fixed and no longer helpful for APT threat actors.
Remember, the situation described in the post only occurred because they used binary images that only a few people could look at. Generating binary kernel on someone’s laptop is easy to subvert in undetectable way, but how do you subvert a Dockerfile stored in Git repo without it being obvious?
"Developer systems are often the weakest link."
(Assuming that the system on itself is designed with security in mind.)
The reason is manifold but include:
- attacks against developer systems are often not or less considered in security planing
- many of the technique you can use to harden a server conflict with development workflows
- there are a lot of tools you likely run on dev systems which add a large (supply chain) attack surface (you can avoid this by allways running everything in a container, including you language server/core of your ides auto completion features).
Some examples:
- docker groub member having pseudo root access
- dev user has sudo rights so key logger can gain root access
- build scripts of more or less any build tool (e.g. npm, maven plugins, etc.)
- locking down code execution on writable hard drives not feasible (or bypassed by python,node,java,bash).
- various selinux options messing up dev or debug tools
- various kernel hardening flags preventing certain debugging tools/approaches
- preventing LD_PRELOAD braking applications and/or test suites
...
A build machine may need to download software dependencies, but ideally those would come from an internal mirror/cache of packages, which should be not just more secure but also quicker and more resilient to network failures.
Interestingly, the way with the least overall headaches is to fully de-privilege all systems humans have access to during regular, non-emergency situations. One of those principles would be that software compiled on a workstation automatically disqualifies from deployment, and no human should even be able to deploy something into a repository the infra can deploy from.
Maybe I should even push container-based builds further and put up a possible project to just destroy and rebuild CI workers every 24 hours. But that will make a lot of build engineers sad.
Do note that "least headaches" does not mean "easy".
And if you increase your defences so much that you're actually somewhat protected from an advanced attacker, you're very, very far on the security vs usability tradeoff, to get there is an organization-wide effort that (unlike simple security basics/best practices) makes doing things more difficult and slows down your business. You do it only if you really have to, which is not the case for most organizations - as we can see from major breaches e.g. SolarWinds, the actual consequences of getting your systems owned are not that large, companies get some bad PR and some costs, but it seems that prevention would cost more and still would be likely to fail against a sufficiently determined attacker.
Do you think having a "kernel guy" building the kernel on his personal laptop before deploying it to production is a very good level of security?
I heard a story from years ago that security researchers tried leaving USB thumb drives in various bank branches to see what would happen. They put autorun scripts on the drives so they would phone home when plugged in. Some 60% of them were plugged in (mostly into bank computers).
I'm not sure when you started using PDFs (I remember mid-90s when my Dad told me about this cool new document format that would standardize formats across platforms, screen and paper!), but hardly anything is static any more.
The author did a good job at making that readable. Is it often like that?
Keep in mind that you actually want your medical provider to have that data, so they can treat you with respect to your medical history, without killing you in the process.
Is medical data really that valuable?
When I was thinking of "defense", I was thinking of the defense contractors who are designing/building things like the next-gen weapons, radar, vehicles, and the like. In that context, when it comes to what they can exfiltrate, I think attackers probably prioritize the details & designs over PII. Just a guess though.
Why not both? Think how valuable the medical information of military staff would be as a source of coersive power.
We normally don't hear about this things not because they can't speak about it but because they don't want to speak about it (bad press).
My guess is that it's a company which takes security relatively serious, but isn't necessary very big.
> hot target [..] else big enough to be a target
I don't thing you need to be that big to be a valid target for a attack of this kind, neither do I think this attack is on a level where "only the most experienced/best hackers" could have pulled it of.
I mean we don't know how the dev laptop was infected but given that it took them 3 month to reinfect it I would say it most likely wasn't a state actor or similar.
HVAC company working in a building where a subcontractor of a major financial firm has an office, for a random example...
It could conceivably belong to a defense organization, but if it did, they wouldn't be able to write up a blog about their findings.
The article just says:
> I wrote a small python program to scan the port 80 traffic capture and create a mapping from each four-tuple TLS connection to a boolean - True for connection with PII and False for all others.
Is it just matching against a list of source IPs? And perhaps the source port, to determine whether it comes from e.g. a network drive (NFS in this case)? Not sure what he uses the full four-tuple for, if this is the answer in the first place. It's very hand-wavy for what is an integral part of finding the intrusion and kind of a holy grail in other situations as well.
https://github.com/capitalone/DataProfiler
Amazon and Microsoft also have their own offerings, but can be quite expensive for network packets (and pretty slow).
Most projects / teams will use some basic regular expressions to capture basics like SSN, credit card numbers or phone numbers. They’re typically just strings of a specific length. More difficult if you’re doing addresses, names, etc.
Love the most recent commit
I saw these regex matchers in school but don't understand them. They go off all day long because one in a dozen numbers match a valid credit card number, even in the lab environment the default setup was clearly unusable. But perhaps more my point: who'd ever upload the stolen data plaintext anyhow? Unencrypted connections have not been the default for stolen data since... the 80s? If your developers are allowed to do rsync/scp/ftps/sftp/https-post/https-git-smart-protocol then so can I, and if they can't do any of the above then they can't do their work. Adding a mitm proxy is, aside from a SPOF waiting to happen, also very easily circumvented. You'd have to reject anything that looks high in entropy (so much for git clone and sending PDFs) and adding a few null bytes to avoid that trigger is also peanuts.
These appliances are snakeoil as far as I've seen. But then I very rarely see our customers use this sort of stuff, and when I do it's usually trivial to circumvent (as I invariably have to to do my work).
Now the repository you linked doesn't use regexes, it uses "a cutting edge pre-trained deep learning model, used to efficiently identify sensitive data". Cool. But I don't see any stats from real world traffic, and I also don't see anyone adding custom python code onto their mitm box to match this against gigabits of traffic. Is this a product that is relevant here, or more of a tech demo that works on example files and could theoretically be adapted? Either way, since it's irrelevant to what the author did, I'm not even sure if this is just spam.
> One thing I didn't get is this magical PII thing. How does the author look at a random network packet -- nay, just packet headers -- and assign a PII:true/false label? I think many corporations would sacrifice the right hand of a sysadmin if that was the way to get this tech.
Checkout Amazon macie or Microsoft presidio or try actually using the library I linked?
It’s usually used in a constrained way, in no way perfect. But it helps investigators track suspected cases of data exfiltration. You can pull something that looks suspect (say a credit card) and compare against an internal dataset and see if it’s legit.
In the repo I linked you can see the citation for an earlier model on synthetic and real world datasets:
https://github.com/capitalone/DataProfiler#references
https://arxiv.org/pdf/2012.09597.pdf
So I don’t really understand the hostility.
What's even more incredible to me is that the researcher somehow recreated exactly the same / correct traffic pattern on their local testing setup, so that they were able to compare the traffic with the production environment to detect that there was a problem. How would you do this?
I'm not even sure what the "time" variable is on the graphs. Response time? (It also seems weird that there's any PII on port 80, but that's an unrelated issue.)
Yeah, that's another thing that has me confused, but I figured one thing at a time...
Thanks for the response, that pre-set PII flag does sound plausible, though it's odd that they'd never mention it and mention a 'four-tuple' instead (sounds like they're trying to use terms not everyone knows? Idk, maybe it's more well-known than it seems to me).
It’s also amazing that they noticed the subtle difference in the NFS packet capture.
I can’t wait for the rest to be published.
Bookmarked
If the attackers have access to brute force OS engineers / sysadmins work pc's then that should probably be the headline. The rest is just about not being found
Reading this, I know of places that have no hope against someone half as decent as this APT. The internet is a scary place
> On March 21, 2021, CNA determined that it sustained a sophisticated cybersecurity attack. The attack caused a network disruption and impacted certain CNA systems, including corporate email. Upon learning of the incident, we immediately engaged a team of third-party forensic experts to investigate and determine the full scope of this incident, which is ongoing.
+ [CNA suffers sophisticated cybersecurity attack](https://www.cna.com/)