You always scale the original image down, it is pretty standard
Also, the reason why we scale down the input image is not for cache effectiveness, it's simply for reducing computations needed. Maybe except for naive matrix multiplication and convolution implementations, which is what this repo does. But there is no point to discuss performance if you are using a naive implementation which by design ignores cache/instruction latency/anything Computer Architecture related. Please, at least take QNNPACK as baseline.