I've always felt that "proper" and responsible LLM use for search [0] would be two boxes: You can describe what you want in the first, and it'll propose search terms in the second, and then those get executed normally.
Yes, the "average user" [1] might not usually care about the second box... Until they need to because the query/results are wrong.
Showing them in tandem means:
1. Users are at least capable of learning through exposure.
2. Users may realize a key term can be added which the model could never have guessed.
3. Users may recognize a term in there that doesn't make sense, allowing them to detect a translation error.
4. If good search-terms leads to a bad outcome, it is possible for someone to report and diagnose it, rather than a fully black-box mystery.
_____
[0] Not just for websites, but also things like internal business software, or SQL queries.
[1] The average that might not exist. ( https://www.thestar.com/news/insight/when-u-s-air-force-disc... ) There are some features that everybody needs, just at different times.
I think the two-tiered input-output approach you propose makes sense. Allowing users to inspect and mutate lower-level languages allows the user to make modifications as needed. I think it's very much in the spirit of free software.
But then again, for text-based search specifically (search engines, notes, etc.), I think there is some value in querying the LLM directly, as it is able to fuzzy search by inspecting its weights. This results in a lesser degree of specificity, which allows for more false positives, but can maybe capture similar words (i.e., synonyms / typos / tenses) or higher-level semantic concepts.
Maybe they're just two different search algorithms, and the user should be able to choose between them.
Thanks for linking the article, I found it very interesting.
And regarding “average doesn’t exist” - that’s true. But no company does a/b testing to land on average. I’d assume a good 85%-kinda pass rate for these type of experiments.
95% of users might not want the double-textbox, but 5% might, and some might want to use it someday.
I think, make enough of these decisions, and there's bound to be mounting friction for users with different preferences navigating your app.
Ideally, the app should offer a way for the user to intuitively configure the app to their own liking.
Ideally, yeah maybe, but why bother with extra implementation, support, costs and etc., when people fold and use the new way anyways?
For a certain cohort (eg those currently between 35 and 45) who did any sort of grade school computer/library classes in the 90's/early2000's, 'chunking' or keyword searching was one of the main skills taught.