A tale of Phobos – How we almost cracked a ransomware using CUDA
cert.pl
cert.pl
Sidecar question: when automating your backups, what's a good way to make sure your rolling backups aren't simply backing up malware-encrypted files? I found out too late that all my backups were encrypted. I only had 1 week retention to save space. Is longer retention and manual checks the only sane strategy? Or has some backup software built in sanity detection for crypto attacks by now?
Compute checksums. Also, if storing diffs check out how many files the back-up think changed. All of your photo library getting re-uploaded should be a red flag.
Backup of Git repositiories:
... # git fsck --full
error: unable to unpack contents of .git/objects/a2/cf1a9631658799733f43c3b3f0a799696a4b21
error: a2cf1a9631658799733f43c3b3f0a799696a4b21: object corrupt or missing: .git/objects/a2/cf1a9631658799733f43c3b3f0a799696a4b21
Oops... No matter if it's a malware, the lack of ECC which by bad luck induced a bit flip that wasn't detected (on an otherwise okay Git repo) or a disk failing, it's trivial to detect if the repo is corrupted.Same for my ripped archive of Audio CDs. The rippers save lots of information and the rips are bitperfect, cross checked with other people's rips' checksums. And the checksums are all there.
For family pictures, I add a checksum to the pictures myself.
Backups aren't really backups until they've been verified :)
So I don't care about the different pictures (or short family movies) format.
I just wrote some Clojure / babashka code to do that. I also truncate the checksum so that the filename doesn't become gigantic: it's not sensitive content, it's just to detect corruption.
Then I can use another computer and generate, say, all the thumbnails of the pictures and do a quick eyeball verification. If it looks correct, later on I can just automatically have the checksums verified.
Funnily enough I got a few old JPG pictures who were corrupt but I ended finding the correct version on older backups.
Checksum then helps too: otherwise you have two files with the same name (say on different HDD), but only one is correct and you don't know which one without manually opening them.
It's not super advanced and maybe a bit overkill but it's not complicated and works fine for my use case.
P.S: I take it another way would be to use a fs that use content-based addressing or does checksumming for me.
Anyone tried using git for 500G of photos?
I would love to if it worked, I have my photo collection spread out on multiple computers and merging the edits to the master backup is always a pita. “Was this file removed from copy A or added to copy B”? All those problems just solve themselves with a clear DVCS git history.
Not having the possibility of ever removing photos, to free up space, is of course another issue of git.
You should check out Git LFS if you want to do that, as it sounds like a good idea in the first place!
If you're open to new tools, git annex is what you're looking for. The other two options are Subversion, which has some DVCS features these days, or Perforce Helix Core (paid), though I can't vouch for it as I've never used it.
Feel free to check it out here https://github.com/Oxen-AI/oxen-release#-oxen would love any feedback!
https://www.scootersoftware.com/v4help/index.html?snapshots....
My backup solution (backuppc/other syncs + zfs + sanoid/syncoid plus offsite server with zfs) means the backup server pulls files from the clients using backuppc/rsync. The backup server volume is zfs snapshoted regularly using sanoid. The offsite server pulls these from the backup server via syncoid/zfs send.
I'm not using rsync.net since I have my own infrastructure, but would definitely choose it as the offsite server if needed.
I can't assist you with recovery, and without a lot of logs and forensics data (or a significant performance improvement), the described method is likely unfeasible. But I'll try to find a matching sample and let you know if it's vulnerable.
> Sidecar question: when automating your backups, what's a good way to make sure your rolling backups aren't simply backing up malware-encrypted files?
Lots of good responses, I like incremental backups without ovewriting anything (supported OOTB by all copy-on-write filesystems, like ZFS or BTRFS). Not sure how to configure this on Windows.
Maybe you could check the level of entropy (measure of randomness) of files before backing up - very high entropy could suggest encrypted data?
Also, JPEG, PNG, .jar, .xlsx, etc. are already compressed, so pretty high entropy to begin with.
As others have pointed out, the growth rate of your de-duplicated backup size is probably the best way to detect ransomware.
I also have most of my non-sensitive data on Onedrive, which keeps old versions of files.
I haven't dive into the code, but only 818000/60=13633 attempts per second for 64 SHA-256 rounds plus one AES-256 decryption on a 2080 doesn't sound right. Learn some GPU and tuning the code can likely increase the throughput a lot.
Mind you, that's only 25x naive Python implementation on (supposedly) 1 core.
Also, hashcat does >1M hash/s for even higher number of SHA256 iterations, and hundreds of million attempts per second on single block AES-256.
Edit: An optimized version of this is very likely deployable on consumer-grade hardware. May still not very useful due to the forensics requirement, though.
There are 256 sha256 iterations on average, so the number is a bit better - but there's probably still a lot to improve (it's much more optimised than the naive version, but it was written by reverse-engineers, not GPGPU specialists). The PoC was also opensourced [0], it would be great if someone experienced could spot any obvious problems.
One major issue was that the key scheduling is basically
do {
key = sha256(key)
while (key[0] != '\0')
So on average there are 256 SHA256s per key, but worst case it takes thousands of iterations. Coding this in a GPU-friendly way was non-trivial, and there are still some GPU cycles wasted.> Edit: An optimized version of this is very likely deployable on consumer-grade hardware. May still not very useful due to the forensics requirement, though.
I think speeding the code three or four orders of magnitude more would go a long way towards pracical usage. I think the biggest issue was getting TID and precise enough time range. Brute-forcing a suspected TID values and widening the time search range would help a lot.
Disclaimer: I worked on that research, but I'm not an employe anymore.
[0]: https://github.com/CERT-Polska/phobos-cuda-decryptor-poc
Now I'm quite fond of python, but largely I read this as the CUDA implementation having quite some room for improvement. Having almost no CUDA experience of my own, that is just a hunch, which I'm glad rfoo supports with (surely) a lot more experience than me.
PIDs on windows are always a multiple of four.
(I didn't see that mentioned in the blog post, but didn't check the code so maybe that's already in there.)
Edit: nvm I guess that's why it says 2^30 not 2^32...
Here's a interesting read that gives a bit of info on the differences and why it doesn't really work. https://rya.nc/asic-cracking.html
Anyone that comes up with a way to repurpose all those bitcoin miners into something productive will be pretty cool in my book.
- hire an expert to find out the missing data (timestamp, is it the vulnerable version?, ...), then
- rent a GPU server and keep the fingers crossed.
I've looked for cheap GPU servers yesterday, the cheapest i found was Ultrarender at 200€ per week for a Dual RTX 3080Ti remote workstation. Which is probably overkill, a single RTX 3080Ti can do it in 33 hours or so (assuming the "12 Nvidia GPUs" they mentioned in the article are of similar speed).
If you want success stories, the same organisation has also published decryptors for other ransomware families (Mapo, Crypromix, Flotera) and they were just broken in a straightforward way. Usually the vulnerability is not explained for successful decryptors, to make bad guys life a bit harder.
phobos is the god and personification of fear and panic in Greek mythology.
this is crucial in understanding ransomware ;)
Perhaps most importantly, also a knight of Mars, beater of ass [2]
1 https://solarsystem.nasa.gov/moons/mars-moons/phobos/in-dept...
More secure data storage at companies where otherwise it would be just silently stolen and sold. More backups. Even some incentive to research security of encryption methods.
We won't get more secure systems without some proper incentives.
Good things can come from bad things. That doesn't make the bad things not bad.
You may prefer world without burglars at all, but that is not an option. It's just wishful thinking.