> This becomes even more dangerous as we're entering the world of LLMs generating SQL[1]
I am beginning to push back against the idea of LLMs authoring SQL. Determinism is really important for most realistic use cases. SQL is one of those "all or nothing" experiences. Pretty much the last thing you want to have a temperature/spiciness slider over.
Granted, if your use case is to simply memorize a handful of canonical queries and to provide a UX-friendly way for business administrators to retrieve these queries, then have at it.
But, if you are hoping to ask the LLM to write novel, sophisticated, domain-specific queries involving joins across 10+ tables, aggregations, recursion, windowing, etc., you are going to have a very bad time. In my experience, this is what everyone is dreaming about doing, and I think it is an unproductive fantasy at this point. Fine tuning will not give you what you seek, but don't let me stop you from trying. We burned over 200 hours on it with nothing to show for our time.
For me, LLMs are far more interesting for determining the user's intentions than emitting a final query. Classification seems like the actual superpower here. Classification can feed a very powerful deterministic query building engine that will actually give you what you want and you can prove it will work. It just takes a lot more tedious work than most humans are willing to endure, so here we are talking about various shortcuts.