14 karma · joined May 16, 2019
However, I did notice that connect took around a minute to a minute and a half before the agent was in the call and able to speak. Is this a byproduct of the underlying calling service you're using or the traffic?
Regardless, awesome app, curious to see how it continues to improve!
On the backend, it uses Stream's Python SDK to capture the WebRTC frames from the player, send them to YOLO to detect their arms and body, and then feed them to the Gemini Live API. Once we have a response from Gemini, the audio output is encoded and sent directly to the call, where the user can hear and respond.
On the backend, it uses the Python AI SDK to capture the WebRTC frames from the player, convert them, and then feed them to the Gemini Live API. Once we have a response from Gemini, the audio output is encoded and sent directly to the call, where the user can hear and respond.
Is anyone else building apps around AI and real-time voice/video? Would be curious to share notes. If anyone is interested in trying for themselves:
Python SDK docs: https://getstream.io/video/docs/python-ai/basics/quickstart/ Github: https://github.com/GetStream/stream-py/tree/webrtc