Just remember there are sensors like the CMV12000 that spew out over 3 billion pixels per second @10bit/pixel and, sufficient light assumed, subpixel motion blur down to iirc about 50% linear overlap or so about. (i.e., only about 50% of the captured pixels have not been captured already.)
Assuming 32 by 32 pixels with a size of 1mm by 1mm, you get a field of view of about 2 by 3 meters or smaller if you want to skip exotic monolithic decoders that integrate de-aliasing into the error correction decoder.
Below about 1.2 sensor pixel per code pixel it stops being fun and you try to find a way to get more magnification.
So, at 150 meters distance, assuming a cube for simplicity, and 5 square meters FOV. We have a cube of 300 m side length. That cube has a surface area of 540 000 m^2, which is just ~100k times the FOV. The 300 frames per second you could get out of the sensor would get this done in about 20 minutes, so yes, it's a lot, but you can reduce to 30% assuming the height difference is approximately known.
Then you can probably get another 10x speed by restricting the angle in which anything interesting could happen, and you get 3% of 20 minutes/1000 seconds. That's half a minute, if you can only tell the scanner that it's about there, with a precision of "between one and two 'o clock, I'm sure".
And that's <10k$ hardware I'm speaking of (not including the cost to get a 2k$ FPGA to extract 2D-barcodes from it's video feed), in 3-digit quantities.
I deem the number plausible.