I'm wondering of how useful pre-assembled programs like "draw detections on an image or video" are. There might be more here, but that's not a very compelling slogan. In my experience, these kinds of tasks only come up when you are demoing something like an object detector, or maybe for diagnostics, but in that case your diagnostics are highly specific to your task at hand. Dataset loaders are useful if you are creating a demo on public datasets, but in product work there is no pre-existing dataset. You're usually extracting the data yourself.
The API for creating polygon zones to filter detections looks nice, but that's a rather limited capability to warrant adopting a whole library. This looks like a toolkit for making computer vision demos, but not something I could see using to build a product. Usually, things like drawing bounding boxes on an image frame, or testing rectangle intersections are the easy stuff.
I do think there's a lot of room for reusable parts in CV, though. Of course the heavyweight in that space is OpenCV, but I wouldn't mind seeing a competitor that doesn't feel like a thinly wrapped C++ library in Python. The multi-view geometry space has few reliable tools in Python, and you spend a lot of time re-implementing classical formulas in NumPy, which I believe could be abstracted away with a 3D geometry toolkit. The kicker is that high-level, readable geometric abstractions (like lens distortion) always end up being a hot path in the code, and at some point they have to get replaced by specialized JIT-ed code.