Thanks! I wonder if the Python number is correct, I remember Python being prohibitively slow in comparison. But assuming it is:
There are 256 sha256 iterations on average, so the number is a bit better - but there's probably still a lot to improve (it's much more optimised than the naive version, but it was written by reverse-engineers, not GPGPU specialists). The PoC was also opensourced [0], it would be great if someone experienced could spot any obvious problems.
One major issue was that the key scheduling is basically
do {
key = sha256(key)
while (key[0] != '\0')
So on average there are 256 SHA256s per key, but worst case it takes thousands of iterations. Coding this in a GPU-friendly way was non-trivial, and there are still some GPU cycles wasted.
> Edit: An optimized version of this is very likely deployable on consumer-grade hardware. May still not very useful due to the forensics requirement, though.
I think speeding the code three or four orders of magnitude more would go a long way towards pracical usage. I think the biggest issue was getting TID and precise enough time range. Brute-forcing a suspected TID values and widening the time search range would help a lot.
Disclaimer: I worked on that research, but I'm not an employe anymore.
[0]: https://github.com/CERT-Polska/phobos-cuda-decryptor-poc