But a lot of the question /responses could be trivial cached. No need to run expensive LLM every time for the same basic "how are you today?" prompts, it only has to be cached once.
Fixed responses for common queries is what we have now.
Not to mention that LLMs tend to be very wordy right now. I’d hate to way 20 seconds to hear my phone say “As a voice assistant I’m not aware of the exact menu of the Thai restaurant on 2nd, but I have opened a google search for it and found the following results.
…”