Show HN: Processing 24 hours of video in ten minutes
sievedata.com
sievedata.com
Sieve is an API that helps you store, process, and automatically search your video data–instantly and efficiently. Just think 10 cameras recording footage at 30 FPS, 24/7. That would be 27 million frames generated in a single day. The videos might be searchable by timestamp, but finding moments of interest is like searching for a needle in a haystack.
We built this visual demo (https://sievedata.com/app/query?api_key=AIzaSyAfKwf0tuuNOHbY...) a little while back which we’d love to get feedback on. It’s ~24 hours of security footage that our API processed in <10 mins and has simple querying and export functionality enabled. The demo is a security use-case but our platform supports a wider variety of use-cases and metadata. We’re working with a few early customers to explore which ones to dive deeper into (happy to hear ideas on this front from the HN community as well!).
To try it on your videos: https://github.com/Sieve-Data/automatic-video-processing
Visual dashboard walkthrough: https://youtu.be/_uyjp_HGZl4
General FAQ: https://sievedata.com/faq
I occasionally download really huge VODs from Twitch or YouTube (mostly gaming tournaments or game playthroughs). I'm quite annoyed of long blocks of breaks, too much blabbering before a match starts, or in the case of bot tournaments -- one of the bots hanging and the system timing out after one whole hour of a mostly static image, etc.
But let me give a more digestible example. The ESL_SC2 channel on Twitch streams StarCraft 2 matches 24/7 (often rebroadcasts). They have these 3 minute-ish breaks with a slightly moving image and an annoying (IMO) music. I believe these segments from the videos can be easily filtered out with applying a particular perception hash on the frames and filtering them out, but never had the time and energy to try it myself.
I'm not asking to be catered to; I'm giving Sieve's authors an idea for a creative endeavor. If they're up for it, download one video from there with youtube-dl / yt-dlp and try it out.
(I suppose this could be useful in the home surveillance scenario as well, e.g. if the camera has been covered with snow for 12 hours, then you don't need the footage and want all frames adhering to that perception hash deleted.)
It works fine at facial recognition for photos but it doesn't even try with videos, any meta data would be useful for hundreds of hours of family videos.
A terabyte of audio and video is only as useful as the time and date we've organised for, and imagine actual usable face ,object recognition for both media types.
Would this work offline for lowend hardware? E.g. Synology 920+
If you have any connects there, feel free to reach out to the email in my bio :)
1) Allow homeowners to take a timelapse of their backyard. For each frame (or some video interval) determine what areas are in sun vs shade. Then greyscale the image and overlay with a heatmap showing hours of sun in each area. Or have 1/2 hour increment bandings.
2) Road monitoring of various types. Traffic count and speed. Common visitors / time of visit (garbage, street cleaning etc, postal, newspaper delivery).
2) This one we currently support + work with. Thanks for bringing it up again :)
Key improvements though would be some type of classification of sun vs not sun for a pixel (based on the entire stream of pixels for the day for that spot to make it easier) because dark stuff in the yard / walls etc messed up my approach. And banding the color (most gardening says things like needs 6 hours of sun a day) so you can clearly see outlines of areas by hours of sun.
Yes, a lot of people don't know how much sun a spot gets and don't want to sit out all day watching (front yard / back yard etc).
I've looked into if there are any sunlight simulators available, but my Google-fu has been lacking. Even if I had a tool that told me "on this date and time, at your location, a 1 meter tall stick casts a shadow _x_ meters in length in _y_ direction", that would be workable.
There are closed form equations that can be used for the angular position of the sun, which will be more than accurate for this purpose.
Then it's mostly some simple trig to determine areas that have a line of sight to the sun.
I have this on my to-do list because in my bedroom on a full moon the light will come through one specific window onto our pillows. i wanted to automate either curtains or something to predict when it'll happen.
Maybe eventually I'll get around to something like this
I'm curious about what projects people in the community are working on that would benefit from this sort of video-frame searching. My computer vision work is in a niche market and I use images, not video. But I figure there must be people building self-driving delivery robots or the like that would like to be able to search their massive video datasets for corner cases where their models need more training data.
Love that you bring up delivery bots. They're the moving example, but you can also think of a parallel application where you have stationary cameras monitoring something of interest (worker safety, defective parts, security, etc).
Interested in hearing where else people are struggling with videos.
I'm currently working with video in the context of gameplay feedback on video games (https://www.volt.school/ is my example app, https://www.vodon.gg/ if you want an instance of your own)
I'd love to be able to detect certain events that happen in the game stream. Things like the player killing someone, picking up a a particular item etc. Adjacent to this, I'd also be interested in your infrastructure around hosting and processing the videos. I'm currently pigging backing off YouTube and in the V2 of my app, I've moved to Cloudflare Stream but I'd love to know how you're working with these massive video streams in a cost effective way.
Our infrastructure is really good at filtering for parts of video where things are happening ("interesting") and being able to parallelize video processing by cutting it up into multiple pieces, running more granular models on those parts, and smoothly interpolating metadata using surrounding frames. You should check out the processing section of our FAQ page (linked in original comment) for a more detailed explanation of how our processing works.
If you're worried about just streaming seamlessly, Mux has a great solution: https://mux.com/live.
Also, for detecting things in e-sports feeds, feel free to reach out to the email in my bio!
Thanks for pointing out the FAQ.
If you're looking at some generic things that are similar to what CLIP was trained on, this would work. Say you're interested in specific physical security metrics, or monitoring defective parts, or specific things about traffic, etc. CLIP might just say "people walking" or "car in intersection" or "part on conveyer belt" which isn't meaningful enough if all your images are exactly that, but with other small differences.
Another important aspect of this is the amount of frames you need to process. Running CLIP even on 27 million frames (1 day of footage) is super expensive. We've built some infra that makes processing video efficient (forms of parallelization + filtering), without you having to think about it.
With your massive parallel infra you're still processing 27 millions frames, right?
> If you're looking at some generic things that are similar to what CLIP was trained on, this would work. Say you're interested in specific physical security metrics, or monitoring defective parts, or specific things about traffic, etc. CLIP might just say "people walking" or "car in intersection" or "part on conveyer belt" which isn't meaningful enough if all your images are exactly that, but with other small differences.
That's exactly transfer learning is for. I am saying to do it with because you automatically get a pretrained model which can do much more than that. Imagine doing a search "a person wearing red cap and yellow hand bag walking toward the exit" or "A person wearing a shirt with mark written on it". Can your system do it right now?
No, it's not. We first run a cheap filter like a motion detector on all of the video, which is inexpensive. We then stack other, more expensive filters on top of this depending on the use-case and eventually run the most expensive metadata-generating models at the end. We also don't do this on every single frame and can interpolate information using surrounding frames. Our parallel infra speeds this up further.
> That's exactly transfer learning is for. I am saying to do it with because you automatically get a pretrained model which can do much more than that. Imagine doing a search "a person wearing red cap and yellow hand bag walking toward the exit" or "A person wearing a shirt with mark written on it". Can your system do it right now?
The issue is that there are very few text-image pair datasets out there, and building a good one is difficult. We constantly use transfer learning in-house when working with different customer data and typical classifier / detector models but haven't yet had success doing so with CLIP. Our system can't semantically search through video just yet but we're exploring the most feasible ways for doing this still. There's some interesting work on this which we've been reading recently:
https://ddkang.github.io/papers/2022/tasti-paper.pdf
https://vcg.ece.ucr.edu/sites/g/files/rcwecm2661/files/2021-...
I am sure the motion sensor is good for a backyard where there's little movement, but not for the front of the house where there's constant movement. I set up motion zones to let me know when there's someone in it, and it triggered every time when someone was caught in the camera, disregarding the zone set. I got so many false alerts each day that I had to turn it off.
But honestly, I want more than what it offers that I cannot tap into... for instance... if someone is trying to break-into my house in the middle of the night I need a phone call or a way connect the camera to my alarm system. A simple phone notification won't wake me up.
It would also be useful to review every time the camera saw a person in the middle of the night when it matters most... not during the day when dozens of people walk in front of it.
I also need to be able to distinguish when it detects a person that's far away, like across the street, and someone who's right in-front of the camera, filling most of the frame.
I'm thinking possibly a live trend analysis / alert sort of thing might be useful if say we detect too much motion, too many people, or some other variable that we determine to be "out of wack".