I think you fundamentally don't understand the topic if you're talking about two pages of text?
The end user would never type in a word of that: they'd say "[Wake word] play me some music"
A piece of software running on a device would transcribe what it heard, and fire off a request to the LLM with all of that text wrapped around their statement.
For ease of sharing I used the web interface to provide the instruction, but you'd use the API with a prompt which also dramatically increases determinism.
No one is writing out the state of each light bulb: you trivially query that information programmatically and bundle it with the request.
—
In a real product there'd be explicit handling of detecting where the request came from, that's already a problem that's been worked on, but I wanted to demonstrate the main difference vs Siri: zero-shot learning
The LLM wasn't told what those volumes mean, but it was flexible enough to infer the intent was to provide a form of location, rather than ask.
It's a forced example so if you want to get caught up on the practicality of audio for locating people be my guest, but it's to show LLMs are great at "lateral applications" of capability:
You give them a few discrete blocks of functionality and limited information, and unlike Siri they can come up with novel arrangements of those blocks to complete a task they haven't yet seen.
—
Honestly the fact you keep going back to "look at all the text" feels a bit like if I showed you the source code for an email messaging app, and you told me: "No one will ever use email! Who would write all that instead of just writing a letter and mailing it?!"