HNHacker News
TopNewBestAskShowJobs

Nash0x7e2

14 karma · joined May 16, 2019

submissionscomments
Nash0x7e2··on The AI Bifurcation of Tech: Why the fundamentals matter more
Been going back and forth about writing this one but the more I thought about it, the more it seemed relevant to today's conversation about Al and growth. In a lot of ways, Al and Agents does allow us to move faster but if the foundation isn't solid (good product experience, developer onboarding, scalability, branding, etc), it almost doesn't matter. Curious if anyone else feels the same.
Nash0x7e2··on Al is killing programming and the Python community
100%, uv is awesome
Nash0x7e2··on Launch HN: April (YC S25) – Voice AI to manage your email and calendar
Looks awesome! I downloaded the app and was able to get it connected to my accounts, and it is working.

However, I did notice that connect took around a minute to a minute and a half before the agent was in the call and able to speak. Is this a byproduct of the underlying calling service you're using or the traffic?

Regardless, awesome app, curious to see how it continues to improve!

Nash0x7e2··on Real-Time Boxing Coaching with Gemini Live, Stream Video and Ultralytics YOLO
Built a demo using Gemini Live and Ultralytic's YOLO models running on Stream's Video API for real-time feedback. In this example, I'm having the LLM provide feedback to the player as they try to improve their form.

On the backend, it uses Stream's Python SDK to capture the WebRTC frames from the player, send them to YOLO to detect their arms and body, and then feed them to the Gemini Live API. Once we have a response from Gemini, the audio output is encoded and sent directly to the call, where the user can hear and respond.

Nash0x7e2··on Gemini Live providing real-time coaching for golf over WebRTC
Built a demo around integrating Gemini Live with Stream's Video API for agent use-cases. In this example, I'm having the LLM provide feedback to players as they try to improve their mini-golf swing.

On the backend, it uses the Python AI SDK to capture the WebRTC frames from the player, convert them, and then feed them to the Gemini Live API. Once we have a response from Gemini, the audio output is encoded and sent directly to the call, where the user can hear and respond.

Is anyone else building apps around AI and real-time voice/video? Would be curious to share notes. If anyone is interested in trying for themselves:

Python SDK docs: https://getstream.io/video/docs/python-ai/basics/quickstart/ Github: https://github.com/GetStream/stream-py/tree/webrtc