While likely, the question asked if there was any improvement shown with other targets to validate that assumption. There is no benefit in thinking.
> And often query results must be 100% accurate and reliable.
It seems that is impossible. Even the human programmers struggle to reliably convert natural language to SQL according to the aforementioned test study. They are slightly better than the known alternatives, but far from perfect. But if another target can get closer to human-level performance, that is significant.
2. The language has to represent a valid computer program. That is as true of SQL as any other target. You can know that it is correct by reading it.
That said, if you have ever used these tools to generate code, you will know that they are much better at some languages than others. In the general case, the target really is the problem sometimes. Does that carry into this particular narrow case? I don't know. What do the comparison results show?