No one. And I mean NO ONE is going to delegate anything like travel or purchasing items to an AI. They just get it wrong too much.
No one. And I mean NO ONE is going to delegate anything like travel or purchasing items to an AI. They just get it wrong too much.
When it comes to money, or really any kind of action that has consequences, you want the assurance of a deterministic outcome. That is simply not possible with this technology. Maybe it will be but I’ve seen nothing to indicate that so far.
OpenAI will subsidize banks to offer that safety-net for users of Operator. If Chase says you're 100% not on the hook for purchase mistakes made by OpenAI's Operator -- are you going to take the risk (as you pointed out, it's financial, it matters to people) with Perplexity Pro or Siri?
Just like with a real assistant, one could set very clear boundaries as to how much the assistant could spend, on what, etc., or even specify that more money can be spent in, say, a particular period and with a few extra passes (using different models?) just to make sure.
It does really feel like the same animal, because to the degree that I remember the discussions, a lot of it was about a perception of some kind of dichotomy (control or not, being able to 'touch' the product or not, etc.) that really doesn't need to exist.
There are cases where this could be true. I could see, say, asking it to do some research for you or something. Or maybe grab some restaurant recommendations. Really anything that brings you options that you can then make a decision on. If I do end up using this tech, thats how I would feel comfortable using it.
When it comes to things like making purchases and such on your behalf, thats where I disagree. Even with boundaries it's just not deterministic enough for most peoples risk appetite IMO. And when I say "most people" I am not talking about HN folks.
And just to be clear, I do use this technology. I just am very aware that its not this magic bullet that the AI bros want us to believe it is. I used it this morning in fact to help me write up some code I felt too lazy to deal with. It very quickly and efficiently wrote what I asked. I was pleasantly surprised and _almost_ missed the very bad bug that it dropped right into the middle of it all. I'm not giving my credit card to that lol.
The fact that such an AI, with sufficient development, even given what we have now, could present everything on a silver platter up to the execution of the /exact/ commands, and this alone is incredibly useful and can save a lot of time/expense/effort.
The crucial bit is that with, say, a human personal assistant, even if they say they will do exactly what you tell them to do, it's still ultimately inexact. But with an AI, it's trivial to 'decorate' a structured 'command object' (as json, xml, whatever) and from that moment let the deterministic system take over.
If anything, we sometimes want an actual human to /not/ deterministically execute /exactly/ what we tell them to, because perhaps they might know we were angry, drunk, or they just heard something on the news that should invalidate our request.
Either way, my point is that this distinction between 'doing research' and 'doing a thing' is not so dichotomous. In practice, I suspect all but the most autonomously-minded people, and especially non-IT folk, given enough sense of control over the confirmation of potentially 'fuzzy' actions, are happy with an AI doing a ton of stuff for them.
I’ve just had an interaction with an insurer who used an ML based system to judge insurability (risk) and make recommendations (requirements) to me for immediate action or policy termination. Fascinating, and totally doable with today’s technology, but it makes me very uneasy.
Anyway, it’s cheaper than hiring an inspector and it gave me some batshit recommendations- including removal of invasive plant material (vines on a trellis, looks great in July) and removing extensive roof debris (leaves, we live in a forest). My recourse is an extremely annoying call center hell, or get on the roof. Guess what I did. And I’m not getting a discount.
That's a bold statement. There are people that get in the backseat or passenger seat while their FSD is driving. There are plenty of other examples of how humans do things because they are reckless or ignorant or any other adjective you want to use
Cursor already reads my private keys and writes code that goes straight to prod. I've stopped validating GPT 4o output for data analysis. Sure, things go wrong from time to time, but the convenience is unmatched.
To delegate our money/critical-wise tasks to an AI (and other things people dream of), first cientists must find a way different technology we have today.
Spoiler: If it ever happens, it's not something we, normal people, will have the pleasure to be fiddling with firsthand. The country capable of discover and develop it first will dominate everything.
(That seems like a lot of trust to me!)
>> Takeover mode: Operator asks the user to take over when inputting sensitive information into the browser, such as login credentials or payment information. When in takeover mode, Operator does not collect or screenshot information entered by the user.
But that assumption is precisely a goal of OpenAI and Anthropic, and it's useful to continue that train of thought to see the bigger picture of where things will go if they get what they want.
The article did a very astute job of that.