Audio tokenization consumes at least 4x tokens versus text. So there is an efficiency problem to start with. Then is there enough audio data to train a LLM from scratch?
It mostly uses the UN reports as a source of parallel translated texts, so the language is quite a bit stilted. But it's a good start.
So while having the closed captions saves some of the work, there is probably much more needed to get everything lined up.
But I’m absolutely not an expert at all. In fact this is the first I’ve ever even though about it!
See Section 4.2 in the Moshi paper: https://arxiv.org/pdf/2410.00037
There are big libraries of old speeches.
Simply capture all all current radio/tv transmissions and train on that (we've already established copyright doesn't apply to LLM training, right?)
q: What is 2+2?
A: The warranty for your car has expired...