79 karma · joined September 5, 2016
Model-Agnostic Infrastructure: Any AI model—open-source, proprietary, LLM, or VLM—can instantly gain real-time vision capabilities. No more waiting for Google to open their doors.
Immediate Availability: Unlike Project Astra, which is still behind Google's beta gates, VideoDB is usable today. Anyone can plug into the API, SDK, or cloud console immediately.
Openness and Developer-Friendliness: Seamless integrations with popular AI tools and frameworks like LlamaIndex, LangChain, and Hugging Face dramatically reduce the barrier to entry. Just a few lines of code and you're live.
The dataset (1,477 manually annotated frames) and benchmarking framework are publicly available to encourage further research.
Paper: https://arxiv.org/abs/2502.06445 Dataset & Repo: https://github.com/video-db/ocr-benchmark
Would love to hear thoughts from the community on the future of VLMs in OCR.
The framework is fully open source and uses a VideoDB key for cloud-based video storage, processing, and streaming. It seamlessly integrates with tools like Stable Diffusion, Eleven Labs, Kling, Replicate, and more.
Looking for collaboration with GenAI audio/ video teams and feedback from amazing devs out here.
Few interesting prompts we tried while building it and loved the results. There's no limit to creativity with this.
*Shark Tank Videos:* [Find every moment where a deal was offered](https://console.videodb.io/player?url=https://stream.videodb...)
*Useful Gadgets* [Show me where the host discusses or reveals the pricing of the gadget](https://console.videodb.io/player?url=https://stream.videodb...)
*Huberman Podcast:* [Find details about every sponsor](https://console.videodb.io/player?url=https://stream.videodb...)
*Masterchef* [Show me the feedback from every judge](https://console.videodb.io/player?url=https://stream.videodb...)
Say goodbye to manual editing, skimming and seeking the video and hello to instant, AI-driven video consumption and creation
But you can use any LLM for analysing the transcript.
RAG applications are great with text, but with video they can't support simple requests like "show me where sleep improvement is discussed"
Here is an example - https://twitter.com/spext_it/status/1286130139290632192
Interactive transcripts are published at https://publish.spext.co