HNHacker News
TopNewBestAskShowJobs

sagarkava

31 karma · joined May 11, 2021

I help companies large and small build the best modern communication experiences with videosdk.live
submissionscomments
sagarkava··on [dead]
Hi HN — I’m launching an open-source WhatsApp AI Voice Agent for phone calls.

Tech stack: It runs on VideoSDK for the SIP gateway, bridging WebRTC ↔ SIP under the hood. For the AI side you can plug in whatever stack you prefer (LLM + STT + TTS). The repo includes example configs.

Why open-source? Most WhatsApp/voice AI projects out there are closed or tied to a single vendor. I wanted something people can actually hack on, fork, and extend — whether that’s experimenting with different voices, building domain-specific agents, or integrating with CRMs.

Performance: End-to-end round-trip latency is ~400–600ms in typical setups. With faster STT/TTS backends there’s headroom to improve this.

I’d love feedback on use cases you’d actually want to build with this: customer support lines, personal AI assistants, language tutors, appointment scheduling, etc. Curious what directions the HN crowd would push this in.

GitHub Repo: https://github.com/videosdk-community/videosdk-whatsapp-ai-c...

Video demo: https://youtu.be/KWfCWE8S_4U?si=yb5WWr4J4n2dgBm8

I’d love feedback: what use cases would you build with this? Customer support, personal AI assistants, language tutors… or something else?

sagarkava··on [dead]
Hi HN, I'm excited to share our new open-source project: an AI voice agent specifically designed for call centers. This project aims to streamline customer interactions and reduce the workload on human agents by automating initial call handling.

Imagine using it to manage customer inquiries, handle reservations, or conduct surveys without human intervention. It's a game-changer for businesses looking to improve efficiency.

Key features include: - Real-time, low-latency voice conversation. - A cascading pipeline using Deepgram for STT, OpenAI (GPT-4o) for LLM, and ElevenLabs for TTS (customizable). - Advanced turn detection and voice activity detection (VAD) for smooth, natural conversations. - Fully open-source and easily customizable. - Support for Agent2Agent and MCP protocols.

Check out the repo: AI Voice Agent for Call Center https://github.com/videosdk-community/ai-voice-agent-for-cal... Main framework: VideoSDK Agents https://github.com/videosdk-live/agents

What use-cases do you envision for this AI voice agent?

sagarkava··on Build an AI telephony agent for inbound and outbound calls
Thanks for the mention. Curious—what challenges are you finding with Pipecat that you're hoping something else (like https://github.com/videosdk-live/agents) might fix?

Always looking to improve based on real gaps devs are facing.

sagarkava··on Build an AI telephony agent for inbound and outbound calls
That is one of our goals. If a solution is under 500 lines of Python, you should not need to pay $499 per month for it. We want to lower the barrier for developers and businesses to build their own voice agents.

GitHub: https://github.com/videosdk-live/agents

https://docs.videosdk.live/ai_agents/voice-agent-quick-start

https://docs.videosdk.live/ai_agents/sip

sagarkava··on Build an AI telephony agent for inbound and outbound calls
The same technology can also enable businesses that never had live phone support to offer it affordably. The goal is augmentation and access, not mass replacement.
sagarkava··on Build an AI telephony agent for inbound and outbound calls
That risk is real. That’s why we made this open-source to empower smaller businesses to build responsible systems with their own logic, prompts, and escalation paths. GitHub: https://github.com/videosdk-live/agents
sagarkava··on Build an AI telephony agent for inbound and outbound calls
That’s valid. But many people, including elderly users, prefer voice interfaces. Our system can serve those customers without requiring a smartphone or web access.

GitHub: https://github.com/videosdk-live/agents

sagarkava··on Build an AI telephony agent for inbound and outbound calls
Only if misused. Our system supports human fallback, logging, and prompt tools to prevent poor user experiences. The key is thoughtful automation. GitHub: https://github.com/videosdk-live/agents Docs on HITL: https://docs.videosdk.live/ai_agents/human-in-the-loop
sagarkava··on Build an AI telephony agent for inbound and outbound calls
Fair point. But when implemented properly, these agents can reliably handle narrow, production-grade tasks like appointment reminders or smart call routing.
sagarkava··on Build an AI telephony agent for inbound and outbound calls
We experienced similar challenges. That’s why we made audio handling, turn detection, and LLM retries modular. You can swap models or providers as needed. GitHub: https://github.com/videosdk-live/agents Blog: https://www.videosdk.live/blog/ai-telephony-agent-inbound-ou...
sagarkava··on Build an AI telephony agent for inbound and outbound calls
That is a fantastic idea. We’ve tested it on low-resource hardware. It’s SIP-agnostic and modular, making it ideal for home or SMB setups. We would love to highlight your RPi implementation if you publish it. GitHub: https://github.com/videosdk-live/agents
sagarkava··on Build an AI telephony agent for inbound and outbound calls
You absolutely can. Our framework can answer calls, run speech-to-text, analyze intent, and respond with LLMs, making it a great defensive tool against spam or scam calls.

GitHub: https://github.com/videosdk-live/agents

sagarkava··on Build an AI telephony agent for inbound and outbound calls
Completely agree. Our goal is to enable smart automation for legitimate, helpful use cases, not intrusive marketing. We also built in Human-in-the-Loop (HITL) support so humans can intervene any time. Docs on HITL: https://docs.videosdk.live/ai_agents/human-in-the-loop GitHub: https://github.com/videosdk-live/agents
sagarkava··on Build an AI telephony agent for inbound and outbound calls
Valid concern. That’s why we focused on responsible use cases such as appointment bots, IVRs, or call reminders, not spam. Our project is open-source, modular, and designed for ethical, contextual automation.

GitHub: https://github.com/videosdk-live/agents

sagarkava··on Build an AI telephony agent for inbound and outbound calls
Totally valid concern. The goal isn’t to replace people, but to offload repetitive tasks so humans can focus on higher-value work.

We’ve also built in Human-in-the-Loop support so a person can step in anytime the AI falls short. More on that here: https://docs.videosdk.live/ai_agents/human-in-the-loop

It’s about shifting roles, not eliminating them and doing it responsibly.

sagarkava··on Build an AI telephony agent for inbound and outbound calls
Hi I am Sagar, We just open-sourced a complete framework to build an AI-powered telephony agent that can handle both inbound and outbound calls—using Python, SIP, and cloud LLMs like OpenAI or Gemini.

You can use it to create smart appointment bots, voice feedback collectors, or even enterprise IVR systems. It’s modular (plug in your SIP provider or AI model), production-ready, and extensible for real-time workflows.

Features include:

SIP & VoIP call handling (Twilio, Plivo, etc)

LLM-integrated AI agent (customizable prompt & tools)

FastAPI-based server for routing and control

Plugins for STT, TTS, sentiment analysis

Support for Agent2Agent and MCP protocols

GitHub Repo:https://github.com/videosdk-live/agents Full Blog: https://www.videosdk.live/blog/ai-telephony-agent-inbound-ou...

Would love feedback from anyone working with telephony, LLMs, or real-time automation!

sagarkava··on Open-source framework for real-time AI voice
Chatterbox is great for local/private TTS with Resemble AI.

voice agent SDK is broader it's full real-time voice infra with STT, LLM, TTS, memory, and RAG built in. You can plug in Resemble, ElevenLabs, etc., and deploy across web, mobile, and telephony with <80ms latency.

sagarkava··on Open-source framework for real-time AI voice
Hey! Quick video overview: https://www.youtube.com/watch?v=m_oc1GDyhrc

Live demo to try it out: https://aiagent.tryvideosdk.live

sagarkava··on Open-source framework for real-time AI voice
Totally fair. The space moves fast, and it's smart to be skeptical. Here's how VideoSDK Real-Time AI Agents stand out from OpenAI agents SDKs and others:

1. Voice infra included OpenAI agents handle logic and memory, but they don’t include real-time audio infra.

VideoSDK gives you:

- <80ms global WebRTC latency

- Built-in turn-taking, VAD, and noise suppression

- Real-time voice across web, mobile, IoT, and telephony

2. Fully modular pipeline No vendor lock-in. Swap STT, LLM, TTS, and avatars. Change models live per user or use case. Want ElevenLabs for tone and OpenAI for reasoning? Easy.

3. Native RAG + memory Integrated long-term memory and retrieval help reduce hallucinations and keep conversations grounded.

4. Scale-ready Deploy globally with one click using Agent Cloud or self-host with full control. Built for production use.

If you're building real-time, voice-first agents that need to work across platforms and scale reliably, this is purpose-built for that.

Happy to dive into your use case if you're exploring options.

sagarkava··on Open-source framework for real-time AI voice
Yes, VideoSDK Real-Time AI Agents are already running in production with several partners across different domains — from healthcare assistants to customer support agents and AI companions. These deployments are handling real user interactions at scale, across web, mobile, and even telephony.

If you're curious about specific use cases or want to explore how it can fit into your product, happy to share more details or walk through an example.

sagarkava··on Open-source framework for real-time AI voice
Hey bigcat12345678, great question!

Yes, with VideoSDK's Real-Time AI Agents, you can control the TTS output tone, either via prompt engineering (if your TTS provider supports it, like ElevenLabs) or by integrating custom models that support tonal control directly. Our modular pipeline architecture makes it easy to plug in providers like ElevenLabs and pass tone/style prompts dynamically per utterance.

We actually support ElevenLabs out of the box. You can check out the integration details here: https://docs.videosdk.live/ai_agents/plugins/tts/eleven-labs

So if you're building AI companions and want them to sound calm, excited, empathetic, etc., you can absolutely prompt for those tones in real time, or even switch voices or tones mid-conversation based on context or user emotion.

Let us know what you're building. Happy to dive deeper into tone control setups or help debug a specific flow!

sagarkava··on Open-source framework for real-time AI voice
Hey

I’m Sagar, co-founder of VideoSDK.

I'm beyond excited to share what we've been building: VideoSDK Real-Time AI Agents. Today, voice is becoming the new UI.

We expect agents to feel human, to understand us, respond instantly, and work seamlessly across web, mobile, and even telephony. But, to achieve this, developers have to stitch together: STT, LLM, TTS, glued with HTTP endpoints and, a prayer.

This most often results in agents that sound robotic, hallucinations and fail in product environments without observability. So we built something to solve that.

Now, we are open sourcing it!

Here’s what it offers:

- Global WebRTC infra with <80ms latency - Native turn detection, VAD, and noise suppression - Modular pipelines for STT, LLM, TTS, avatars, and - real-time model switching - Built-in RAG + memory for grounding and hallucination resistance - SDKs for web, mobile, Unity, IoT, and telephony — no glue code needed - Agent Cloud to scale infinitely with one-click deployments — or self-host with full control Think of it like moving from a walkie-talkie to a modern network tower that handles thousands of calls.

VideoSDK gives you the infrastructure to build voice agents that actually work in the real world, at scale.

I'd love your thoughts and questions! Happy to dive deep into architecture, use cases, or crazy edge cases you've been struggling with.

sagarkava··on Domino that called exit(); for Twilio's Programmable Video
Twilio has decided to shut down its Programmable Video SDK. This decision, while understandable, may leave many developers wondering about the future of their video-based applications.

In a recent statement, Twilio CEO Jeff Lawson explained the reason behind this decision: While I was disappointed with this decision, the Twilio Video SDK was one of the best products in town, especially for builders.

sagarkava··on [dead]
We made it to#1 Product of the Day on Product Hunt

And, we're excited that Video SDK 2.0 is Trending on the #3 Product of the Week

It would be really helpful if you can share a Video SDK 2.0 with your FluttreFlow community.

Support us HARD, and make Video SDK the #1 Product of the Week

https://www.producthunt.com/posts/video-sdk-2-0

sagarkava··on [dead]
I just want to thank you for your support & love!

We made it to #1 Product of the Day on Product Hunt with your help.

And, we're excited that Video SDK is trending on the #3 Product of the Week

It's time for you to show your love again, support us HARD, and make Video SDK the #1 Product of the Week

https://www.producthunt.com/posts/video-sdk-2-0

Help us turn this launch into a success party by spreading the word https://ctt.ac/Dvysi

sagarkava··on [dead]

   Build an Interactive Live Streaming App in Flutter
  Learn how to build an interactive live streaming app in flutter using Video SDK with Yash Chudasma, SDE at Video SDK on May 31st, 2022 at 6:00 PM IST. 
What's in it for you?

Learn how to interactive live streaming work

Learn how to deploy & integrate flutter video SDK (Android & iOS)

Learn how to add features to the flutter SDK - recordings, interactive chat, HD screen sharing, etc

sagarkava··on [dead]
We created interactive Witeboard to be the "fastest" way to share ideas remotely.

It's the fastest way to collaborate real-time with your team.

Explain Everything interactive Whiteboard is a browser and cloud-based platform That delivers a truly cross-platform collaborative whiteboarding experience.

https://docs.videosdk.live/docs/guide/prebuilt-video-and-aud...

sagarkava··on [dead]
Thank you for share this blog sir.
sagarkava··on [dead]
Hey Hello! I am Sagar, the CMO of Videosdk.live. Super elated, I am making this announcement for you all. We have come up with something exciting for the developer community.

Videosdk.live has recently launched its low code video SDK for video conferencing on the web and mobile and it turns out to be a huge success. Believe me, it is amazing and effortless. This prebuilt sets a video conference in your application in no more than 10 minutes, or precisely in just one click.

𝗜𝘀 𝘁𝗵𝗮𝘁 𝗽𝗼𝘀𝘀𝗶𝗯𝗹𝗲? Yes! Undoubtedly yes! We have made it possible. I believe smart work is today a more wanting deal than hard work for achieving successful tasks in a shorter span. We have made that provision for you. Access your meeting smartly. All the hard work is on us, you just need to add the spark.

We build products that prosper effective communication and provide a range of services all in one video call platform. Build video calls for applications and websites hassle-free! Add these amazing features along:

𝗣𝗮𝗿𝘁𝗶𝗰𝗶𝗽𝗮𝗻𝘁𝘀 𝗚𝗿𝗶𝗱: The participants’ grid is an intelligent grid system to make the participants more visible. Now the participants and host can see each other and derive active communication in the best quality. The video quality automatically adjusts based on the grid size to optimize internet usage.

𝗥𝗮𝗶𝘀𝗲 𝗵𝗮𝗻𝗱: Solve the problem of waiting in queue to get your queries acknowledged. Ask a question or share your opinion with just a one-click button. The Raise hand button helps the participants to get along with their doubts without disturbing the speaker in the middle of his speech. Happens to be effortless right? That’s how we maintain decorum.

𝗦𝗰𝗿𝗲𝗲𝗻 𝘀𝗵𝗮𝗿𝗲: Share your screen and present anything at the highest resolution. Cast a window, chrome tabs, or an entire screen in a few seconds. Trust me, it is a no lag feature, and it wouldn’t get stuck in any manner. Nonetheless, screen sharing won’t bother the audio and video quality of your presentation.

𝗣𝗮𝗿𝘁𝗶𝗰𝗶𝗽𝗮𝗻𝘁𝘀 𝗹𝗶𝘀𝘁: Now you can look for all your participants in one click. Search through the crowd of 1000+ in the same meeting. Recognize the presence of any participant at the meeting, without any long search efforts. Mute and unmute the participants, allowing them to speak in the meeting.

𝗣𝘂𝗯𝗹𝗶𝗰 𝗰𝗵𝗮𝘁: Make valuable conclusions in a large crowd through chats. Broadcast a message or start a discussion on the spot. The Public chat window enables people to communicate with each other and derive suggestions and ideas within the meeting. No need for discussions over other applications. Use Public Chat and let’s not struggle with multiple apps.

𝗥𝗧𝗠𝗣 𝗟𝗶𝘃𝗲𝘀𝘁𝗿𝗲𝗮𝗺: Host an event and stream the meeting to YouTube, Facebook, and other RTMP services stress-free. Make big announcements at an ease, visible to a huge crowd from several social media in one go. Serve your conference to millions of people and make them aware of your brand and the decisions you’ve taken for its growth.

𝗖𝗹𝗼𝘂𝗱 𝗿𝗲𝗰𝗼𝗿𝗱𝗶𝗻𝗴: Watch later and share the meeting with those who couldn’t join. What can be a better deal when the whole meeting can be played over multiple times after it has happened? You can now play meetings on-demand with cloud recording and share the access to people to make it appreciated by the non-attendees.

𝗪𝗵𝗶𝘁𝗲𝗯𝗼𝗮𝗿𝗱: This flawless feature allows to draw, write & scribble to explain the most complex ideas effortlessly. Create amazing ideas on the whiteboard and share them with your participants to gather focus on discussions more accurately.

𝗗𝗲𝗰𝗶𝘀𝗶𝗼𝗻 𝗽𝗼𝗹𝗹: Allow agendas to come to a mutual vote even if the meeting has a huge participant crowd. Ask for opinions or take a survey and arrive at a quick decision. Make tasks effortless with the poll. Derive on mutual opinion in the presence of all participants. Make decisions worthwhile. https://www.videosdk.live/

sagarkava··on Ask HN: How do you build your personal brand?
Write articles and post them on LinkedIn, Medium, or a personal website. Create a personal website to showcase your portfolio. Print business cards with a tagline that highlights the thing you’re known for. Incorporate your tagline into your elevator pitch. Seek out small speaking engagements.
Page 1 of 2Next →