Why do you need LLM to interpret patterns?
Maybe I'm not understanding it, but as I get it, LLMs weren't really important: all they did was further interpreting outputs of a fronting audio-to-text classifier model.
You don't need them, but they are one way to do it that people know how to implement.
Identifying patterns is fairly amenable to analytic approaches, interpreting them, less so.