The Articulation Barrier: Prompt-Driven AI UX Hurts Usability
uxtigers.com
uxtigers.com
To notice that people can't use a natural language interface because they're bad at organizing and expressing their thoughts, then claim this is a failing of the interface is for the poor craftsmen to blame their tools.
To say that we should make language models that are somehow agnostic to a user's command of language is so obviously a fools errand its hard to know where to begin. The author is implicitly calling for mind-reading models.
Or in other words, are UX designers catering to an audience that ought to not exist in the first place.
No, its not.
The UI is the output of the craft, not the tool of the craftsman. And the blame isn't coming from the craftsman, anyway.
It is just as true as the one about craftsmen, if not as popular a saying, that it is a poor toolmaker who blames the people who the tools are made for the tools' unsuitability for their purpose. If a user interface doesn't work for the actual users, it is a poor UI.
You might say that the toolmaker's job is to make a foolproof tool that does its job completely independently of the user's skill. But my third line denies this possibility in the context of language models.
No this is the equivalent of a hammer that only drives one very specific brand and size of nail. Language is a notoriously terrible tool.
It is one of the best available. It's the default, built-in method of communication for humans.
Admittedly without numbers to bear this out, I've been consistently impressed by the quality of modern text-to-speech audio transcription, and LLMs are, when provided with written dialogue, fairly skilled at replying with usable-to-spectacular results. This suggests a serious potential for improved usability for these low literacy users by expressing their queries as a verbal and untyped stream-of-consciousness by using transcription, while still giving them back well-formed results in response.
Additionally, to the extent that this is a problem for non-native speakers separate from native speakers specifically, one of the benefits of LLMs is that they often are capable of generating well-formed inputs in multiple languages trained in their corpus - this suggests that even on an interface designed without localization in mind, Prompt-driven UX can potentially give you user-language localization essentially for free.
Given these two points, I don't like the assertion that these kinds of interfaces a priori hurt usability, because they assume a framework of interaction that first doesn't have to be used, and second ignores the users whose usability would potentially be greatly empowered by these kinds of interfaces.
The truth is, these LLMs have already seen the average (and worse) of human writing, and have generalized to a more precise (and typically the intended) input. When they're wrong, they recover pretty well when corrected. The existence of prompt engineers is more for the adaptation and additional context for the mass of thin apps on top of an LLM.
I posit that what's needed isn't an abandoning of LLMs as interface, but an acknowledgement that chat-based AI as an input only helps at a certain, rather high, level of abstraction. There are many possible applications that can be built with an LLM but very few that actually enhance the UI of the rest of the app. Application designers need deeper understanding of the LLM's capabilities and the core value of the software product it's intending to enhance in order to build something that is an improvement on what the app would be without an LLM.
For much of the value of LLMs, prompts absolutely do not have to be well written. For some things they do; when I want it to do some programming I'm much clearer. Or if I'm trying to get it to do something it's very bad at (reasoning), but mostly I avoid that.
There's a reason why Prompt Engineer is now a job title: companies need someone able to take someone's half-baked ideas and massage them into something coherent enough to elicit a useful response.
It certainly gives me pause for thought - while LLMs right now are worse than useless due to hallucinations meaning you cannot trust a single thing they say, I am sure that will improve over the years. But if people cannot even express or articulate what they want (even well-educated, highly-literate people) then I can't see LLMs replacing too many jobs in the near to mid term.
Yet if an essential quality to using LLMs is dealing with hallucinations your ability to detect them really limits your ability to use the tool. Some people believe the IRS wants them to buy Google Play gift cards from Walmart and redeem them.
This is absolutely not true. You should try Bing copilot (or whatever they are calling it this week). It provides extensive citations for everything it asserts. You can easily verify facts by following the hyperlinks. I agree that the original Chatgpt interface is very hard to trust on factual matters due to the lack of citations.
I think this supports the article quite well too. This idea that LLMs 'hallucinate' or that you can't 'trust' them is a clear illustration of the underlying problem.
I've never seen an LLM hallucinate or give an untrustworthy response. But I've seen many many people claim that they do.
Those people are usually attempting to use a language model to do mathematical calculations, instead of using a calculator. Or to lookup facts, instead of using a database. Or to get an overview of a topic, instead of using an encyclopedia.
So naturally, they have about as much luck as someone trying to chop down a tree with a phillips head screwdriver or trying to drive a screw into the wall with a chainsaw. Or even, trying to use a phillips head screwdriver to drive a phillips head screw, but doing so by holding the tip and banging the handle against the screw as if it were a hammer.
Too many people aren't using the right tool for the job, and even if they are, they don't know how to use it correctly. That requires some understanding of the tool, some aptitude, and in the case of a language model, some literacy.
But these LLMs are being marketed as if they are general purpose artificial intelligences that can do anything, when they're obviously not. And their users don't understand that and aren't using them properly.
However, if you use them to pseudorandomly generate natural sounding text, given their training and context, then they absolutely do in fact do that correctly. No 'hallucinations', no 'wrong answers'. They randomly generate natural enough text which fits the topic, pretty much every time.
The issue is marketing them as something they're not, to people who aren't literate enough on the topic to know the difference.
And if, as the article claims, half the people aren't even literate enough to use them, let alone understand when and how to use them, then it's not a good idea to use them as a standard interface nor to make promises about what they can do, when the people that would be using them wouldn't be able to make them do those things.
Back in the day, you pretty much had to be able to understand and program a computer to use one. But imagine if your standard cellphone or laptop nowadays just booted up to a BASIC prompt or even a DOS prompt. Most people would be lost. These days we have UIs such that almost anyone can use them, without that understanding and without having to write their own programs.
LLMs are still in that earlier stage. Where you need to understand them to get value out of them. But they're being dolled up and advertised as if you don't.
I agree with you, they can't successfully replace jobs when people don't even know what they want or how to use them.
I think we overestimate just how much people like to browse for stuff rather than have it auto-magically done for them or shown to them. When it comes to online shopping, you probably prefer to browse, unless it's a recurring purchase. I still use a search engine for code issues because I prefer browsing for answers vs having one generated for me, usually because I like the additional context that comes from browsing.
Outside of browsing, I find it really bizarre that all the generative AI art tools out there use chat interfaces rather than a suite of generative graphics tools, i.e. a generative paintbrush, but then again, most people using said tools aren't trying to create some kind of specific thing in their minds eye, they just want an image. Otherwise, if you wanted to paint a picture, then you'd want either a paintbrush or some kind of paintbrush analogue. Painting a picture through a chat interface just seems kind of ungainly.
Kind of an extreme example, but imagine trying to drive a car that has all the familiar controls replaced with a chat interface. Even the A/C is better served by a dial of some sort rather than asking the AI to make the interior a little bit cooler.
Seriously, what is the point they're trying to make here?
Accessible design is also often mandated by legislation which has existed for many decades before something like ChatGPT existed and old laws have to be interpreted in light of the new technology coming out.
(Maybe even users rapidly learning to express themselves better and more concisely, as a more general skill. Or maybe not: maybe tools like LLMs will teach "bad" habits, like just spamming tokens at a gist of vague intent, and not caring much about the quality, or doing it in a rapid refining loop, without pausing to think and articulate clearly.)
Possible precedent: Very early in Google history (it might've still been at Stanford?) a famous HCI person came by my office, and I had Google on my screen. On a tangent, he started editorializing (with the irritation of someone who cares), on how their affordance is just a text box, and people don't know what to type there.
My thought at the time (as a mere student) was that the Web search UI was already familiar (Google's was only a bit more all-in on the text box than previously; they mostly just dispensed with the conventional noise on the page), and users and Google would both figure it out.
And both did (Google, especially, with smarter responses, based on what they'd learned of intent and information over time). That worked out pretty well for everyone, for a long time.
Like how many people communicate now? :(
You click the magic button and behind the scenes it automatically adds a bunch of fancy tokens, like "iridescent" or "affluent", that trigger the desired effect.
It also makes me think of a few people who I should try to convince to get into prompting…
Especially because for now the LLMs aren't very good with the native languages of most of the world's population. https://news.ycombinator.com/item?id=39130990