HNHacker News
TopNewBestAskShowJobs

ashu_trv

79 karma · joined September 5, 2016

Building @videodb_io I like building fundamental systems. Wanders in deep thoughts of technology, science, spirituality and human nature.
submissionscomments
ashu_trv··on GPT-5: Overdue, overhyped and underwhelming. And that's not the worst of it
I agree. I am big fan of o3, and GPT 5 is not the same, it's like going back to GPT-3 level stupidity. It doesn't care about context, feels super dumb.
ashu_trv··on Live video feed for every multimodal model not just Gemini
What's really impressive here:

Model-Agnostic Infrastructure: Any AI model—open-source, proprietary, LLM, or VLM—can instantly gain real-time vision capabilities. No more waiting for Google to open their doors.

Immediate Availability: Unlike Project Astra, which is still behind Google's beta gates, VideoDB is usable today. Anyone can plug into the API, SDK, or cloud console immediately.

Openness and Developer-Friendliness: Seamless integrations with popular AI tools and frameworks like LlamaIndex, LangChain, and Hugging Face dramatically reduce the barrier to entry. Just a few lines of code and you're live.

ashu_trv··on Auto-Sync Your Docs, SDKs and Examples for LLMs and AI Agents
Yeah, we tried to solve the 1 and 3. 2nd is still an open problem. Can you share more about the MECE?
ashu_trv··on Auto-Sync Your Docs, SDKs and Examples for LLMs and AI Agents
Keeping documentation and SDK updates aligned with evolving "LLM contexts" can quickly overwhelm dev teams. At VideoDB, we've built an open-source solution—Agent Toolkit—that automates syncing your docs, SDK versions, and examples, making your dev content effortlessly consumable by Cursor, Claude AI, and other agents. Ready-to-use template available.
ashu_trv··on Vision-Language Models vs. Traditional OCR in Video – New Benchmark
A new benchmark study evaluates Vision-Language Models (Claude-3, Gemini-1.5, GPT-4o) against traditional OCR tools (EasyOCR, RapidOCR) for extracting text from videos. The findings show VLMs outperforming OCR in many cases but also highlight challenges like hallucinated text and handling occluded/stylized fonts.

The dataset (1,477 manually annotated frames) and benchmarking framework are publicly available to encourage further research.

Paper: https://arxiv.org/abs/2502.06445 Dataset & Repo: https://github.com/video-db/ocr-benchmark

Would love to hear thoughts from the community on the future of VLMs in OCR.

ashu_trv··on Show HN:Video is hard: until now
Thanks! Added now.
ashu_trv··on Show HN:Video is hard: until now
This open-source agent framework is like ChatGPT, but for videos. It simplifies complex video tasks like search, editing, compilation, and—best of all—generation. The results stream instantly. You can even extend the agents to suit your needs and build custom automated workflows.

The framework is fully open source and uses a VideoDB key for cloud-based video storage, processing, and streaming. It seamlessly integrates with tools like Stable Diffusion, Eleven Labs, Kling, Replicate, and more.

Looking for collaboration with GenAI audio/ video teams and feedback from amazing devs out here.

ashu_trv··on Show HN: Instantly create video clips from LLM prompts
It analyse the transcript, but there is no way to get back the video clip without building your own video infra. We at Videodb are solving the exact problem.
ashu_trv··on Show HN: Instantly create video clips from LLM prompts
LLMs are great with text, but they don't help you consume or create video clips. Checkout PromptClip - Use natural language to describe the what you want. - Instantly get video clips with the help of LLMs like OpenAI or Claude.

Few interesting prompts we tried while building it and loved the results. There's no limit to creativity with this.

*Shark Tank Videos:* [Find every moment where a deal was offered](https://console.videodb.io/player?url=https://stream.videodb...)

*Useful Gadgets* [Show me where the host discusses or reveals the pricing of the gadget](https://console.videodb.io/player?url=https://stream.videodb...)

*Huberman Podcast:* [Find details about every sponsor](https://console.videodb.io/player?url=https://stream.videodb...)

*Masterchef* [Show me the feedback from every judge](https://console.videodb.io/player?url=https://stream.videodb...)

Say goodbye to manual editing, skimming and seeking the video and hello to instant, AI-driven video consumption and creation

ashu_trv··on Show HN: Instantly create video clips from LLM prompts
Yeah, VideoDB is the next-gen infrastructure for videos and actually less costly than current video infrastructure.

But you can use any LLM for analysing the transcript.

ashu_trv··on Show HN: GPT-Powered Video Retrieval and Streaming
Build custom GPT on your video data with StreamRAG in 2 mins.This search agent find relevant moments across hundreds of hours of content and return a video clip instantly.

RAG applications are great with text, but with video they can't support simple requests like "show me where sleep improvement is discussed"

ashu_trv··on Show HN: Twitter bot generates interactive transcript of any audio/video
Made a twitter bot that lets you generate an interactive transcript(Spext Docs) of any audio/video in a tweet

Here is an example - https://twitter.com/spext_it/status/1286130139290632192

Interactive transcripts are published at https://publish.spext.co

ashu_trv··on DevSpace – The Fastest Developer Tooling for Kubernetes Development
Very cool! How is it going to be beneficial over the native solutions provided by cloud providers itself?