This C code isn't doing exactly the same thing; here it's calculating a global similarity measure by calculating sum-of-squared-differences and comparing the end sum against the specified threshold, instead of comparing each pixel-wise diff against the threshold.
It is straightforward C code but it's vectorised by the compiler: https://godbolt.org/z/-3uO1z
Excluding the overhead of reading & decoding the PNGs, it takes 80ms for the same 8400x4725 images as my parent comment.
Benchmark:
$ git clone git@github.com:stb-tester/stb-tester.git
$ cd stb-tester
$ make
$ ipython
>>> import _stbt.sqdiff, cv2, timeit
>>> f1 = cv2.imread("water-4k.png")
>>> f2 = cv2.imread("water-4k-2.png")
>>> timeit.timeit(lambda: _stbt.sqdiff._sqdiff_c(f1, f2), number=1)
0.081
[1] The Python code in my parent comment will allocate memory (the size of the entire image) for each intermediate calculation like the colourspace conversion, the absolute differences, etc.