Congrats on the launch! What are you using to make the LLM understand a video file?
Are you doing transcription + sending frames to a vision or is there a third party service for this?
Are you doing transcription + sending frames to a vision or is there a third party service for this?