This is super hard. Not only do you have hidden functionality (what a great insight!) but they are trying to do something better than the current state of the art (which is human-to-human). What do I mean? Just listen to yourself when you get something over the phone. You don't call the restaurant and say "give me a reservation for four at 8." You typically use an interactive process of query and response ("can you fit four in at 8pm?" "Yes, today." "No, not five, four". Then they repeat it back to you just in case.
We get annoyed at how crappy the voice recognition systems are, and they are crappy, but human voice recognition reminds me of a TCP session negotiation, or maybe the baud rate negotiation of a bell-compatible Hayes-style modem. It's hardly a one-shot process, yet we want our machines to be.