410 karma · joined September 8, 2019
mokshith@sievedata.com
https://www.sievedata.com/
The model tries to copy the blinks of the original video so it's possible that in other conditions, you'd notice less of this.
Fun to see this feedback though, definitely something worth improving :)
https://www.sievedata.com/blog/eye-contact-correction-gaze-c...
Newer models have come out that allow the same thing to be done and control even more than the eyes.
See here: https://github.com/KwaiVGI/LivePortrait/blob/main/assets/doc...
For web-conferencing, local use is great so NVIDIA's tools are what we recommend in that case.
Would recommend trying it on other videos, it is surprisingly good. Although there definitely are areas to improve.
* Process videos up to 12x faster than realtime * Costs <$0.01 / min of video * Combines visual and audial components
The goal here is not to build a single E2E model but something that could actually be used in production while preserving relatively high quality.
You can try it out yourself here: https://www.sievedata.com/functions/sieve/describe How we built it: https://www.sievedata.com/blog/describe-video-summary-beta-l... The code: https://github.com/sieve-community/describe
1. We process videos, and can store their metadata / embeddings in perpetuity. This means a simple query API, don't worry about storing a DB of metadata yourself.
2. "Reverse Image Search" - Say a user cares about something specific, like a blue box that's toppled over. Using high-dimensional representations of certain frames in video, we can surface every time this scenario shows up in the future, based on a user having marked it as "interesting" in the past.
3. Our "killer" app is the video processing infra that's super modular. We give this flexibility to the full-stack dev to plug in their own models, or use ours.
A final thing to note is that we're going after a completely different set of companies, specifically companies focused on building software for a specific industry (i.e. https://www.equipmentshare.com/) that eventually want to add video analytics as an offering. It's also unclear who Azure or GCP are targeting though it seems like it's for general media content.
Most use-cases have limitations in bandwidth, so not all data is streamed up to the cloud. Typically, people will run simple motion detectors on the physical camera device itself, and save content in the cloud that they think might be valuable. Our service can then be used to make that content searchable, past simple motion detection.
There are other use-cases where people store tons of data in the cloud already. Media content is an example of this. In this case, companies might want to analyze all the raw data for which we then charge only for intervals of video where there's real action. If the video is static, we won't charge the full per minute pricing.
1. People have definitely tried this in the past in terms of having similar tech, but I don't think anyone has sanely approached it at the same angle. We're seeing so many more companies like EquipmentShare (https://www.equipmentshare.com/), UpKeep (https://www.upkeep.com/), Spot AI (https://www.spot.ai/) and others which are building core software for industries without necessarily having tons of video expertise in-house. Also other video-first companies like Gong (https://www.gong.io/), Mux (https://mux.com/), and Loom (https://www.loom.com/) are building platforms that might benefit from visual analysis at scale. Things like serverless functions were also less cool before.
2. There's a bunch of ways to search using Sieve. Some are fixed variables, others are visual similarity based. See here: https://sievedata.com/faq
3. Our target is the full stack developer within companies building cloud-first software for specific industries. Easy-to-use APIs are key because of this. They are the ones with the expertise to package it into a product that really works for the end user, and we're the ones with computer-vision and video expertise.
We currently process on the per-frame basis but have the optionality to increase the window size of 1, to any arbitrary size like 300 frames. This is what allows us to currently detect "actions" for some of our customers. Of course, it's not the same as a sliding window but it's something we're considering.
No, it's not. We first run a cheap filter like a motion detector on all of the video, which is inexpensive. We then stack other, more expensive filters on top of this depending on the use-case and eventually run the most expensive metadata-generating models at the end. We also don't do this on every single frame and can interpolate information using surrounding frames. Our parallel infra speeds this up further.
> That's exactly transfer learning is for. I am saying to do it with because you automatically get a pretrained model which can do much more than that. Imagine doing a search "a person wearing red cap and yellow hand bag walking toward the exit" or "A person wearing a shirt with mark written on it". Can your system do it right now?
The issue is that there are very few text-image pair datasets out there, and building a good one is difficult. We constantly use transfer learning in-house when working with different customer data and typical classifier / detector models but haven't yet had success doing so with CLIP. Our system can't semantically search through video just yet but we're exploring the most feasible ways for doing this still. There's some interesting work on this which we've been reading recently:
https://ddkang.github.io/papers/2022/tasti-paper.pdf
https://vcg.ece.ucr.edu/sites/g/files/rcwecm2661/files/2021-...
If you have any connects there, feel free to reach out to the email in my bio :)
Our infrastructure is really good at filtering for parts of video where things are happening ("interesting") and being able to parallelize video processing by cutting it up into multiple pieces, running more granular models on those parts, and smoothly interpolating metadata using surrounding frames. You should check out the processing section of our FAQ page (linked in original comment) for a more detailed explanation of how our processing works.
If you're worried about just streaming seamlessly, Mux has a great solution: https://mux.com/live.
Also, for detecting things in e-sports feeds, feel free to reach out to the email in my bio!
2) This one we currently support + work with. Thanks for bringing it up again :)
If you're looking at some generic things that are similar to what CLIP was trained on, this would work. Say you're interested in specific physical security metrics, or monitoring defective parts, or specific things about traffic, etc. CLIP might just say "people walking" or "car in intersection" or "part on conveyer belt" which isn't meaningful enough if all your images are exactly that, but with other small differences.
Another important aspect of this is the amount of frames you need to process. Running CLIP even on 27 million frames (1 day of footage) is super expensive. We've built some infra that makes processing video efficient (forms of parallelization + filtering), without you having to think about it.
I'm thinking possibly a live trend analysis / alert sort of thing might be useful if say we detect too much motion, too many people, or some other variable that we determine to be "out of wack".
Love that you bring up delivery bots. They're the moving example, but you can also think of a parallel application where you have stationary cameras monitoring something of interest (worker safety, defective parts, security, etc).
Interested in hearing where else people are struggling with videos.
Sieve is an API that helps you store, process, and automatically search your video data–instantly and efficiently. Just think 10 cameras recording footage at 30 FPS, 24/7. That would be 27 million frames generated in a single day. The videos might be searchable by timestamp, but finding moments of interest is like searching for a needle in a haystack.
We built this visual demo (https://sievedata.com/app/query?api_key=AIzaSyAfKwf0tuuNOHbY...) a little while back which we’d love to get feedback on. It’s ~24 hours of security footage that our API processed in <10 mins and has simple querying and export functionality enabled. The demo is a security use-case but our platform supports a wider variety of use-cases and metadata. We’re working with a few early customers to explore which ones to dive deeper into (happy to hear ideas on this front from the HN community as well!).
To try it on your videos: https://github.com/Sieve-Data/automatic-video-processing
Visual dashboard walkthrough: https://youtu.be/_uyjp_HGZl4
General FAQ: https://sievedata.com/faq
I'm actually the author of the post as well. Two things: 1. I stated 10 cameras, not 1 so according to your calculations that would mean 190 GB per day instead of 19 GB. 2. I didn't make the assumption that there is always a unique frame but rather got this estimate by trying to save videos using OpenCV across various types of cameras. The variable I didn't change was the environment I took these videos in, which tended to be high activity ones where many frames did have some motion in them. 1 TB might be an overshoot sometimes, but the point of adding it into the article was to highlight the order of magnitude for a setup, that all things considered is pretty small. Another thing is that in the context of computer vision training data, the number of frames is something we care about more than the amount of space they take up which is why I also state 27 million frames.
Thanks for pointing it out though. Might be better not to even mention video file size given the high variability.