The trick was just to let the LLM come up with its own SQL queries for searching... and the results are impressive.
The trick was just to let the LLM come up with its own SQL queries for searching... and the results are impressive.
>On this benchmark, a pure LLM generated an accuracy score of zero. Adding RAG, prompt engineering, and agentic AI raised accuracy to the 10+% range.
[1] Any text-to-SQL benchmark should address difficulties of real-world data stores (acm.org) (21 comments):
https://news.ycombinator.com/item?id=49013995
[2] If You Think You Can Do Real-World Text-to-SQL:
https://cacm.acm.org/blogcacm/if-you-think-you-can-do-real-w...
[3] BEAVER: An Enterprise Benchmark for Text-to-SQL:
Before a lot of frameworks existed, you'd see DEVs taking user input on a web form, and then just throwing it directly at the MTA. So spammers could submit email@address\nCC: persontospam@address, and the like.
Now LLMs are a different beast, but you have input validation for LLMs, unique to all other validation methods. Yet there's actually no safe way to ever validate user input for a LLM, except for very rigid input validation on single words. Take the email example above. You'd need a regex to only validate an email address (and that isn't simple), but once you expand it to actually allowing sentences?
The LLM is now input validation vulnerable.
And that means no user input can be used in unvalidated commands.
And then just random hallucinations. I'm curious how the gp managed weirdo LLM behaviour, like out of the blue 'drop table' or accidental select into as opposed to just select.