Yeah, you really shouldn’t be building this (or really any sort of important or destructive interface) with LLMs unless you have an intermediate representation which users can inspect and understand. Even if the understanding of English were perfect there are still ambiguities you’d need to sort out. You also need to be able to suggest something useful if the user asks a nonsense question (for example data that doesn’t exist in your schema or its contents). When I worked on this a decade ago we were quite good at mapping unknown queries to various suggestions drawn from a canonical subset of English for which we had very explicit semantics. But obviously the Soave of unknowns was huge.
I’ve no doubt the AIs will reach human parity at some point (and therefore still make mistakes!) but I’d be terrified deploying any sort of one-shot text-to-results pipeline today.