Very impressive, kudos. I'd like screenshots rather than PSNR and SSIM because those don't translate into human perception of a good encode.
The only grip I have is the 13k lines of C + intrinsics.
The only grip I have is the 13k lines of C + intrinsics.
How are you supposed to get acceptable performance without intrinsics?
I also wonder how much the compiler can do autovectorisation on code like this --- it's pretty much exactly the type of code that autovectorisation is intended for.
Edit: I noticed in the benchmark that it compressed the 10s foreman.cif (demo video) in half a second, so it's already 20x faster than realtime on that small resolution.