I'm streaming ~20 fps (17 to 30) 720P directly from my home IP4 address, and when a person is in-frame long enough and caught by the tracker, a stream goes to an AWS endpoint for storage.
I've experimented with both SSDMobileNet and Yolo3, which are both pretty error prone but they do a much better job filtering out moving tree limbs and passing clouds, unlike Arlo.
You need way more processing power than an RPi to do this at 30fps, and C/C++, not Python. (There are literally dozens of projects for the RPi and TFlow online but they all get like 0.1 fps or less by using Flask and browser reload of a PNG... great for POC but not for real video)
I wrote very little of the code, honestly: only the capture pipe required a new C element. I started with NVidia DeepStream which is phenomenally well-written, and their built-in accelerated RTSP element, and added a custom GStreamer element that outputs a downsampled MPEG capture to the cloud when the upstream detector tracks an object. NVidia also wrote the tracker, you just need to provide an object detector like SSDMobileNet or YOLO. NVidia gets it.
The main 4 camera-pipe mux splits into the AI engine and into a tee to the RTSP server on one side and my capture element on the other side.
It was amazingly simple, and If I turn the CCD cameras down to 720P with h265 and a low bitrate, I don't need to turn on the noisy Xavier fan. The onboard Arm core does the detected downsampling (one camera only, a limitation right now) and pushes the video with a rest endpoint on a node server in the AWS cloud.
I'm very pleased with it, I haven't tested scaling but if I turned off the GPU governors I could easily go to 8 cameras. I went with PoE because WiFi can't handle the demand.