Please could you give an overview of how this actually works? Have some ideas of where the tech could be useful but not sure how I'd actually go about implementing it. Do you have a GPT model on a server and code to transcribe the video then summarise the transcription. Or do you use one of the APIs from OpenAI?
If you use their APIs:
* How costly has it been to run your service? (If you don't mind answering)
* Is it customisable? If you wanted to run a chat bot for example, would you be able to make it understand the request (I'd assume something similar to an 'intent' when developing Alexa skills) and give it data so it knows the answer?