I have about 48 different cameras where I want to count people and get their approximate location in the frame.
I want to run an object detection model on all of those video streams simultaneously.
My AWS instance maxes out after 7 simultaneous streams so I figured I don't really need real-time monitoring. One frame every couple of seconds, even every minute could potentially suffice, since I am dealing with larger time-frames. Since I don't want to run too many instances at the same time, what are some viable strategies to achieve this?
My plan is to have 5-6 instances of the ML model loaded up and waiting to accept a frame. When one of them is ready, it will instruct one of the RTSP streams to send it a frame, which it will process and store / send the result to an application server. I feel like I may not even be able to consume so many RTSP streams at once (I've never tried so I don't know), so I may have to have some other method of priming the handshake etc. before the model asks for a frame to process.
Is there a better / non-hacky way of achieving this (i.e. managing the workload on a single GPU instance) ?
I don't have any control of the camera hardware at all.