No, I expect people to stop pretending that LLMs somehow know context unlike the stupid Siri.
There's considerably more in that prompt besides just "you need to act like a home assistant"
There was insufficient context.
Imagine I tell you "turn on that light, where I'm pointing". You'd do no better. No one here is under the conviction magical prescience is involved. This tooling provides the mechanism for an initial API call to be tied to the event described, in natural language, as "look where I'm pointing".
The first response (to ask for clarification) is precisely what a human agent would do to get context to clarify the coarse-grained request. The second guess, assuming you disabled the (explicit) allowance for clarifying questions, is also a magnificent recognition of implicit, common-sense context. Seems it's even more effective than you at following the true context for this tools appropriate placement.