Denying interoperability is so culturally ingrained at this point, that it got pretty much baked into entire web stack. The only force currently countering this is accessibility - screen readers are pretty much an interoperability backdoor with legal backing in some situations, so not every company gets to ignore it.
No, we'll have to settle for "chat agents" powered by multimodal LLMs working as general-purpose web scrappers, because those models are the ultimate form of adversarial interoperability, and chat agents are the cheapest, least-effort way to let users operate them.
For example, McDonald's has heavily shifted away from cashiers taking orders and instead is using the kiosks to have customers order. The downside of this is 1) it's incredibly unsanitary and 2) customers are so goddamn slow at tapping on that god awful screen. An AI agent could actually take orders with surprisingly good accuracy.
Now, whether we want that in the world is a whole different debate.
Today, you need to bypass "do you have the app", "do you want fries with that", "do you want to donate", "are you sure you don't want fries?" and a couple more.
All this is exactly what your parent comment was saying: "To a business, the user interface isn't there to help the user achieve their goals - it's a platform for milking the users as much as possible."
Regarding sanitation, not sure if they are any worse than, say, door handles.
What I wrote earlier, about business seeing the interface as a platform for milking users, applies just as much to human interface. After all, "do you want fries with that?" didn't originate with the Kiosks, but with human cashiers. Human stuff is, too, being programmed by corporate to upsell you shit. They have explicit instructions for it, and regular compliance checks by "mystery shoppers".
Now, the upsell capabilities of human cashier interface are limited by training, compliance and controls, all of which are both expensive and unreliable processes; additionally, customers are able to skip some of the upsells by refusing the offer quickly and angrily enough - trying to force cashiers to upsell anyway breaks too many social and cultural expectations to be feasible. Meanwhile, programming a Kiosk is free on the margin - you get zero-time training (and retraining) and 100% compliance, and the customer has no control. You can say "stop asking me about fries" to a Kiosk all day, and it won't stop.
It's highly likely a voice chat interface will combine the worst of the characteristics above. It's still software like Kiosk, just programmed by prompts, so still free on the margin, compliant, and retrainable on the spot. At the same time, the voice/conversational aspect makes us perceive the system more like a person, making us more susceptible to upsells, while at the same time, denying us agency and control, because it's still a computer and can easily be made to keep asking you about fries, with no way for you to make it shut up.
It will depend on the material of the door handles. In my experience, many of the handles are some kind of metal, and bacteria has a harder time surviving on metal surfaces. Compare that to a screen that requires some pretty hard presses in order to get registered inputs from it, and I think you'd find a considerably higher amount of bacteria sitting there.
Additionally, I try to use to my sleeve in order to open door handles whenever possible.
Note - because this is something which needs to be pointed out in any discussion of AI now - even though human beings also make mistakes this is still markedly less accurate than the average human employee.
(I say model, but for this problem I'd consider a pipeline where the powerful model is just parsing orders and formulating replies, while being sanity-checked by a cheaper model and some old-school logic to detect excessive amounts or unusual combinations. I'd also consider using "open source" model in place of GPT-4o, as open models allow doing "alignment" shenanigans in the latent space, instead of just in the prompts.)
About as unsanitary as opening the door to get to the kiosks. The kiosks get wiped down more than the door.
LLMs, if used at all, aren't aware enough to even know what the software can do, and many actual chat UIs are worse than that!
My "favourite" design pattern for chat UIs is to invite you type, freely, whatever you like, then immediately enter a wizard "flow" like it's 1991 and entirely discard everything you typed. Pure hostility.
> it's incredibly unsanitary
I never thought about this. Does McD's PR team have anything to say about it? I assume that a bunch of people have challenged them about it on Twitter or TikTok. Would you feel better if there was a kind of automatic/robotic window washer that sanitised the screen after each use?The key to me about the kiosks is: (1) initially, replace cashier labour costs with new expensive machines, and (2) medium-to-long term, upgrade the software with more and more "upsell" logic. This could be incredibly effective as a sales tactic. (Not withstanding the possibility, I fully agree with your final sentence!)
Can you imagine if a celebrity, like Kim Kardashian or David Beckham, lent their likeness for a fee to McD's to create an assistant that would talk with you during your order? (Surely, AI/ML can generate video/anime that looks/moves/sounds just like them.) I can foresee it, and it would be the near-perfect economic exploitation of parasocial relationships in a retail setting.
They probably ignore them, as they should - the same problem exists everywhere, from ATMs to door keypads to stores to self-checkout to tapping your card on stuff, etc.
> initially, replace cashier labour costs with new expensive machines,
Labor, like energy, is conserved in the system. It might be easier to counter proliferation of those systems if the narration was focused less about companies replacing labor on their side, and more on the fact that this labor gets transferred to the customers, who are now laboring for free for the company, doing the same things that used to be done better and faster by a dedicated employee.
The irony is, the reason neither Siri nor Alexa nor Google Assistant/Now/${whatever they call it these days} nor Cortana achieved this isn't the voice side of the equation. That one sucks too, when you realize that 20 years ago Microsoft Speech API could do better, fully locally, on cheap consumer hardware, but the real problem is the integration approach. Doing interop by agreements between vendors only ever led to commercial entities exposing minimal, trivial functionality of their services, which were activated by voice commands in the form of "{Brand Wake word}, {verb} {Brand 1} to {verb} {Brand 2}" etc.
This is not an ergonomic user interface, it's merely making people constantly read ads themselves. "Okay Google, play some Taylor Swift on Spotify" is literally three brand ads in eight words you just spoke out loud.
No, all the magical voice experience you describe is enabled[0] by having multimodal LLMs that can be sicced on any website and beat it into submission, whether the website vendor likes it or not. Hopefully they won't screw it up (again[1]) trying to commercialize it by offering third parties control over what LLMs can do. If, in this new reality, I have to utter the word "Spotify" to have my phone start playing music, this is going to be a double regression relative to MS Speech API in the mid 2000s.
--
[0] - Actually, it was possible ever since OpenAI added function calling, which was like over a good year ago - if you exposed stuff you care about as functions on your own. As it is, currently the smartphone voice assistant that's closest to Star Trek experience is actually free and easy to set up - it's Home Assistant with its mobile app (for the phone assistant side) and server-side integrations (mostly, but not limited to, IoT hardware).
[1] - Like OpenAI did with "GPTs". They've tried to package a system prompt and function call configuration into a digital product and build a marketplace around it. This delayed their release of the functionality to the official ChatGPT app/website for about half a year, leading to an absurd situation where, for those 6+ months, anyone with API access could use a much better implementation of "GPTs" via third-party frontends like TypingMind.
Dynamically generated interactive UIs are something people are barely beginning to experiment with; we don't know if current models can do them reliably for realistic problems, and how effort has to go into setting them up for any particular product. At this point, they're an expensive, conceptual solution, that doesn't scale.
And would this company spend billions of dollars for this infinitesimally small increase in convenience? No, of course not; you are not the real customer here. Consider reading between the lines and thinking about what you are sacrificing just for the sake of minor convenience.
"I stamp the envelope and mail it in a mailbox in front of the post office, and I go home. And I’ve had a hell of a good time. And I tell you, we are here on Earth to fart around, and don’t let anybody tell you any different...How beautiful it is to get up and go do something."
Dont be limited with these examples.
How about Airline booking, try different airlines, go to the confirmation screen. then the user can check if everything is allright and if the user wants to finish the booking on the most cheapest one.
What's the goal of technology ? Automate everything so that we don't have to live anymore ? We might as well build matrix pods at that point
Talking to an x-Model (still not AI), just like talking to a human, has never been, is not now, and will never be faster than looking at an information-dense table of data.
x-Models (will never be AI) will eat the world though, long after the dream of talking to a computer to reserve a table has died, because they are so good at flooding social media with bullshit to facilitate the sales of drop-shipped garbage to hundreds of millions of people.
That being said, it is highly likely that is an extremely large group of people who are so braindead that they need a robot to click through TripAdvisor links for them to create a boring, sterile, assembly-line one-day tour of Rome.
Whether or not those people have enough money to be extracted from them to make running such a service profitable remains to be seen.
The Rome trip is even more absurd. Part of the fun of a trip is figuring out what you want to do.
This seems like a product aimed at the delusional, self important, managerial class.
Ok but that does not mean others share the same opinion. Try doing a walk in for a fancy restaurant on the weekend, see how that goes?
But… what if I told you that AI could generate an context-specific user interface on the fly to accomplish the specific task at hand. This way we don't have to deal with the random (and often hostile) user interfaces from random websites but still enjoy the convenience of it. I think this will be the future.
Now with the power of AI we have added back in that middle man to countless more services!
If these are your pain points in life, and they're worth spending $500b to solve, you must live in an insane bubble.
Maybe it could read HN for me and tell me if there is anything I'd find interesting. But then how would I slack off?
I am not looking forward to a trip booked for wrong dates with the hotel name confused/hallucinated for a different one.
Let the bot deal with the ads, the cookie banners, the upsells, "newsletters" and all of the other web BS we deal with.
The bot clicks through the front door of the website, just like us. No APIs, no keys, no nothing.
"Hey Siri, grab me a bottle of slow release 500mg Vitamin C from either Amazon or Walmart, whichever has the best deal. Kthx"
and I'm surprised that people don't bring this visualisation up more often.