You've picked a poor example, I think. And you're probably coding in an inefficient way.
Here's the OpenCV way for a pair of 2048x2048 images:
import cv2
import time
t = time.clock()
a = cv2.imread("./image_1.tiff", cv2.IMREAD_GRAYSCALE);
b = cv2.imread("./image_0.tiff", cv2.IMREAD_GRAYSCALE);
c = a-b
cv2.imwrite("out.tiff", c);
print time.clock()-t
Takes about 0.2 CPU seconds on average for me (note using time.clock, not time.time on UNIX).
#include <opencv2/opencv.hpp>
#include <ctime>
using namespace cv;
using namespace std;
int main(void){
clock_t begin = clock();
Mat a = imread("./image_1.tiff", IMREAD_GRAYSCALE);
Mat b = imread("./image_0.tiff", IMREAD_GRAYSCALE);
Mat c = a - b;
imwrite("out.tiff", c);
clock_t end = clock();
double elapsed_secs = double(end - begin) / CLOCKS_PER_SEC;
cout << elapsed_secs << endl;
}
Again, about 0.2 seconds. The difference is negligible if you use the right libraries. Python should not be your bottleneck for high performance code.