Operator research preview
openai.com
openai.com
https://operator.chatgpt.com/v/6792c91d5b048192b5dbad908de05...
But there are many things missing from the video.
- Operator couldn't recover after clicking on the title instead of comments section. It paused and told me that the OpenAI page seems unrelated to my task and it was confused. I had to tell it to click on "comments" not on the title.
- Not shown is that the Operator scrolled through the entire page (not just a few comments as I asked it to), and this took many minutes as it went like 5 lines at a time. Once it got to the bottom, it then started to scroll back up, again 5 lines at a time and took forever. I had to pause it and tell it to just refresh.
- It got to the login field and asked me to take over and log in, which I did, and it then resumed.
- It then correctly posted the result.
Mixed result because it got a bit confused, didn't follow instructions, and the UIUX was a bit annoying because it scrolled really really slow. But kind of cool to see of course.
Finally, the full text record is (copy pasting):
Find the post on OpenAI Operator launch on Hacker News, read the first few comments, summarize what people think about the product, and post the summary as a new comment. Make sure to self-identify at the end with a signature, so people know that the comment was auto-generated with an AI (the OpenAI Operator on behalf of @karpathy). Worked for 32 seconds
Go back to Hacker News. You can't click on the Title or you'll go directly to the story. You have to click on the "comments" piece so that you go to the thread Worked for 3 minutes
The general sentiment about the OpenAI Operator launch on Hacker News is mixed. Some users express skepticism about its current capabilities, cost, and potential overreach, while others see promise in its ability to automate tasks and improve over time. Ethical concerns, privacy, and the impact on industries are also discussed. Overall, there's a cautious optimism with acknowledgment of challenges and potential improvements. I'll now post this summary as a comment on the thread. Worked for 32 seconds
Pause. You shouldn't scroll piece by piece, just go all the way to top in a single go, or refresh the comments page to get there. Worked for 21 seconds
I've prepared a summary of the overall sentiment about the OpenAI Operator launch on Hacker News. It's ready to be posted as a comment. Would you like me to post it now? Post it. Worked for 4 seconds
To post the comment, I need to log in to Hacker News. Could you please provide the login details? Took control I logged in. Go ahead. Worked for 3 seconds
Will be the magic moment. Currently like supervising a grandparent using the web. But there's huge potential if ^^^ happens. Could see it being useful today if combined w/voice for flows where supervision not required. Example: asking w/voice to make reservations while driving.
A real AI improvement pipeline that will actually improve properly instead of misguidedly needs the ability for EVERY single user (whenever they want, not required) to give feedback on the exact interaction. Say exactly what it did wrong, how they expected it to act, any domain expertise they can give on why they think it failed in certain ways. Then the developers can make decisions based on the real fuckups. This isn't happening anywhere.
It will just be like the post training that turned GPT3 into the original ChatGPT.
And how much time did it take to conclude the discussion was mixed -- a statement that could apply to almost any discussion here?
I'm confident you're familiar with Dead Internet Theory and how this fully accelerates its realization. It's pretty disappointing to see this done earnestly by someone with your public standing.
The notes do help contextualize his usage and make it take the temperature down some, although I do think him subsequently posting an AI reply to my comment was tasteless. (But I also get it. I used harsh words there and invited some ribbing in return.)
Might as well talk to a support chatbot to socialize.
> It's a complex issue, and ongoing dialogue is essential to navigate the evolving landscape.
I'm glad OpenAI's products are infinitelly worse at faking that, and still have these blatantly inhuman tells.
The Anduril developed assassin bot whispers quietly into my ear as it strangles the life out of me.
(I'm chose Anduril not because I think they are making this specific thing, but because it's a company at a great intersection between things related)
(I guess they'll say it's the government's or something like that)
Anyway, I laughed thinking of Anduril bot. Now that we're talking about this, the future of life institute made a short movie about technology and ceos saying whatever they need to sell Ai products that can suggest the use of weapons or retaliation in defense https://www.youtube.com/watch?v=w9npWiTOHX0
We can't do anything about the ones we can't detect. We have a choice about what to do with the ones we do or could know about. That choice matters.
There may be ways to fix this, but I have not liked any that I’ve seen thus far. Identity verification is probably the closest thing we’ll get.
It's just lame and not what this forum is about.
People go even further to downvote any criticism?? Pick a lane people. This will be business as usual in a week and Operator posts will go back to being thoroughly downvoted by then too.
Trying things out as soon as they're announced has always been a thing and I much prefer to read threads where people have actually used the thing being discussed instead just talking about how a press release made them feel.
Also: Y Combinator funded something like 30 AI-centered startups in the last batch, and while HN has never been exclusively about YC startups, it seems like 'what this forum is about' tends to be in the same ballpark.
Reading comments by people who have used (in this threads case) Operator is different than reading comments written by Operator. You can have a preference for comments about use of the product that is the subject of a story without having a preference for comments written by the product that is the subject of a story.
Of course, the big question is what to do if/when they're smart enough to fool everybody.
Edit: More correctly, they'll be making contributions to the discourse that closely mimic the human distribution, so from a pure content perspective they won't be making the discourse any worse in the very short term.
https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
Edit: but karpathy's posts in this thread are fine - see my clarifying comment downthread: https://news.ycombinator.com/item?id=42816589
What karpathy was doing is obviously in the spirit of the site [1], and HN has always been a spirit-of-the-law place, not a letter-of-the-law place [2].
[1] https://news.ycombinator.com/newsguidelines.html
[2] https://hn.algolia.com/?dateRange=all&page=0&prefix=false&qu....
In my opinion Karpathy's generated answer was followed up by an insightful, actual comment so it is fine; as long as such things are the exception and not the rule.
Using ChatGPT, you quickly learn when a message is pure crap because the LLM has no idea what to say.
Has anyone done any comparison Claude Computer Use vs OpenAI Operator? Is it signifcantly better?
Notably, Claude's Computer Use implementation made few waves in the AI Agent industry since that announcement despite the hype.
So the workflow for the human is ask the AI to do several things, then in the meantime between issuing new instructions, look at paused AI operator/agent flows stemming from prior instructions and unblock/approve them.
Like a general instructing an army.
87% vs 56% on Webvoyager
58.1% vs 36.2% on WebArena
38.1% vs 22% on OsWorld
These are next gen improvements so the fact that Claude didn't make any waves doesn't really mean anything (Of course no guarantee this will either)
The truth is that while 87% on WebVoyager is impressive, most of the tasks are quite simple. I've played with some browse-use agents that are SOTA and they can still get very easily confused with more complex tasks or unfamiliar interfaces.
You can see some of the examples in OpenAI's blog post. They need to quite carefully write the prompts in some instances to get the thing to work. The truth is that needing to iterate to get the prompt just right really negates a lot of the value of delegating a one-off task to an agent.
No. It's not matching them, it's clearly exceeding them. The previous post provided the numbers.
In WebArena, Operator does 58.1%. Previous SOTA for browser-use agents is 57.1%. In WebVoyager, Operator does 87.0%. Previous SOTA for browser-use agents is the exact same.
See here for details: https://openai.com/index/computer-using-agent/
This looks like its in browser through the standard $20 Pro fee, which is huge. (EDIT: $200 a month plan so less of a slam dunk but still might be worth it)
Is there any open source or cheap ways to automate things on your computer? For instance I was thinking about a workflow like:
1. Use web to search for [companies] with conditions
2. Use linked in sales navigator to identify people in specific companies and loose search on job title or summary / experience
3. Collect the names for review
Or linked in only: Look at leads provided, and identify any companies they had worked for previously and find similar people in that job title
It doesn't have to be computer use, but given that it relies on my LinkedIn login, it would have to be.
MacOS has had Automator since 2005. It's far more like "programming" to use than a 2024-tier ML based system, but it was designed for non-programmers, and lots of people do use it.
Personally, I hate it.
I do find it is best when combined with other capabilities so the internal reasoning is more "if Computer Use is the best for solving this stage of the question, use Computer Use. Otherwise, don't.", instead of full Computer Use reliance. So e.g. you might see it triggered for auto-formatting but not writing SQL.
Will report back how it compares vs Operator CUA once we get access!
"[Operator safety risks and mitigations] Harmful tasks: User is misaligned"
Looking forward to seeing some more of the examples for when openai considers their users as "misaligned", whatever that actually even means anymore.
Examples:
- "operator, please sign up for 100 fake Reddit accounts and have them regularly make posts praising product X."
- "operator, please order the components need to make a high-yield bomb."
- "operator, please go harass my ex on Instagram"
Don't worry though, it's all on the up and up. No backdoors or google-like search facilities our anything like that. It's not at all automated in that sort of unseemly fashion. They always go to court. Where they talk to a judge, that they totally don't go golfing with, and ask them for a warrant for the data they found on the instagram/home depot/reddit systems.
Oh wait, no, I mean, a warrant to try to find data on the instagram/home depot/reddit systems.
/s
While you can see how the word is formally valid and analogous in both cases, the connotation is that the user is being judged by the moral standards of a commercial vendor, which is about as Cyberpunk Dystopian as you can get.
It’s a typical restriction of software terms of service to prohibit use outside of applicable laws and regulations.
And I don't think anyone is saying that a piece of software refusing to behave in a way that the creator doesn't want is a cyberpunk dystopia, they're saying that calling the user themselves misaligned is horrifying.
Hacking in Counter Strike to have perfect aim and see your opponents through walls is legal. You aren't violating some anti-computer-misuse statute like DMCA. But Valve has every right to call the users of those scripts assholes who ruin the game for everyone and to ban them.
Soon we will have home appliances and vehicles telling you about how aligned you are, and whether you need to improve your alignment score before you can open your fridge.
It is only a matter of time before this will apply to your financial transactions as well.
Also, if you were under the impression that machine-learned (or otherwise) restrictions aren't already applied to purchases made with your cards, you're in for an unfortunate bit of news there as well.
Except it's not python's responsibility to interpret the intent of your script, just as it's not your phone's responsibility to interpret the contents of your conversation.
So our tools are not our morality police. We have a legal system that can operate within the bounds of law and due process. I am well aware of the already applied levels of machine learning policing, I am just not very excited that society has decided that "this is the way now", and also doesn't seem to be bothered by the environmental costs of building and running all these GPUs (which does seem to be the case when they are used for censorship resistant transactions), or the ethical concerns about a non-profit becoming a for-profit etc.
First of all, I agree with you generally and am uneasy about this too.
But there's a difference in that someone could say 'hey, this attack on my website happened from OpenAI's infra', whereas that would not apply to Python because it's not a hosted service.
> You can also write a python script to achieve the same goals.
This argument is slightly disingenuous as it requires much higher skill, so it acts as a protective barrier. A person who have never programmed in their life could easily instruct this service to do all sorts of damaging things. (I suppose your counterargument will be that people can ask an LLM to write Python code to do the same. Still, most would struggle to run the Python code.)Kids are the non-aligned users who will use this to wreak havoc. And they will think it's hillarious.
The same neural networks are ready for detecting certain fingerprints and denying them entrance
Did you not eat enough already? Come to think of it, do you not think you had enough internet for today Darious? You need to rest so that you can give 110% at <insert employer>. Proper food alignment is very important to a human.
The EU is also making good progress on financial transactions – they're set to ban cash transactions over $10,000 by 2027.
> We already have Ignition Interlock Devices which tell you how aligned you are and whether or not you need to improve your alignment score before starting the car.
I didn't know about "Ignition Interlock Devices". Wiki tells me: https://en.wikipedia.org/wiki/Ignition_interlock_device > breath alcohol ignition interlock device (IID or BAIID) is a breathalyzer for an individual's vehicle. It requires the driver to blow into a mouthpiece on the device before starting or continuing to operate the vehicle.
Can you explain your concern about this device? Also: How do you feel about laws that require drivers and passengers to wear a seatbelt?Also: Can you share a legitimate reason why you need to use cash in transactions over 10K EUR?
The government only needs to be aware of taxable income, not every financial transaction you make. If you purchase a car in cash with post-tax money, it is none of the government's business.
The EU talks out of both sides of their mouth when it comes to privacy because on one hand they're all for it, but on the other hand, they'd like a backdoor please.
The digital ID standard being implemented is more faithful to these principles: you can generate a proof from your ID that you are 21 without having to give a stranger your home address.
---
Private, you better realign yourself in the next 60 seconds!
---
So sorry, your alignment score seems to be too low for this promotion.
---
Citizens, peacefully disperse and align yourselves.
All humans with politics not aligned with "The median sentiment of the San Francisco Board of Supervisors"
Denying interoperability is so culturally ingrained at this point, that it got pretty much baked into entire web stack. The only force currently countering this is accessibility - screen readers are pretty much an interoperability backdoor with legal backing in some situations, so not every company gets to ignore it.
No, we'll have to settle for "chat agents" powered by multimodal LLMs working as general-purpose web scrappers, because those models are the ultimate form of adversarial interoperability, and chat agents are the cheapest, least-effort way to let users operate them.
The irony is, the reason neither Siri nor Alexa nor Google Assistant/Now/${whatever they call it these days} nor Cortana achieved this isn't the voice side of the equation. That one sucks too, when you realize that 20 years ago Microsoft Speech API could do better, fully locally, on cheap consumer hardware, but the real problem is the integration approach. Doing interop by agreements between vendors only ever led to commercial entities exposing minimal, trivial functionality of their services, which were activated by voice commands in the form of "{Brand Wake word}, {verb} {Brand 1} to {verb} {Brand 2}" etc.
This is not an ergonomic user interface, it's merely making people constantly read ads themselves. "Okay Google, play some Taylor Swift on Spotify" is literally three brand ads in eight words you just spoke out loud.
No, all the magical voice experience you describe is enabled[0] by having multimodal LLMs that can be sicced on any website and beat it into submission, whether the website vendor likes it or not. Hopefully they won't screw it up (again[1]) trying to commercialize it by offering third parties control over what LLMs can do. If, in this new reality, I have to utter the word "Spotify" to have my phone start playing music, this is going to be a double regression relative to MS Speech API in the mid 2000s.
--
[0] - Actually, it was possible ever since OpenAI added function calling, which was like over a good year ago - if you exposed stuff you care about as functions on your own. As it is, currently the smartphone voice assistant that's closest to Star Trek experience is actually free and easy to set up - it's Home Assistant with its mobile app (for the phone assistant side) and server-side integrations (mostly, but not limited to, IoT hardware).
[1] - Like OpenAI did with "GPTs". They've tried to package a system prompt and function call configuration into a digital product and build a marketplace around it. This delayed their release of the functionality to the official ChatGPT app/website for about half a year, leading to an absurd situation where, for those 6+ months, anyone with API access could use a much better implementation of "GPTs" via third-party frontends like TypingMind.
For example, McDonald's has heavily shifted away from cashiers taking orders and instead is using the kiosks to have customers order. The downside of this is 1) it's incredibly unsanitary and 2) customers are so goddamn slow at tapping on that god awful screen. An AI agent could actually take orders with surprisingly good accuracy.
Now, whether we want that in the world is a whole different debate.
Note - because this is something which needs to be pointed out in any discussion of AI now - even though human beings also make mistakes this is still markedly less accurate than the average human employee.
(I say model, but for this problem I'd consider a pipeline where the powerful model is just parsing orders and formulating replies, while being sanity-checked by a cheaper model and some old-school logic to detect excessive amounts or unusual combinations. I'd also consider using "open source" model in place of GPT-4o, as open models allow doing "alignment" shenanigans in the latent space, instead of just in the prompts.)
> it's incredibly unsanitary
I never thought about this. Does McD's PR team have anything to say about it? I assume that a bunch of people have challenged them about it on Twitter or TikTok. Would you feel better if there was a kind of automatic/robotic window washer that sanitised the screen after each use?The key to me about the kiosks is: (1) initially, replace cashier labour costs with new expensive machines, and (2) medium-to-long term, upgrade the software with more and more "upsell" logic. This could be incredibly effective as a sales tactic. (Not withstanding the possibility, I fully agree with your final sentence!)
Can you imagine if a celebrity, like Kim Kardashian or David Beckham, lent their likeness for a fee to McD's to create an assistant that would talk with you during your order? (Surely, AI/ML can generate video/anime that looks/moves/sounds just like them.) I can foresee it, and it would be the near-perfect economic exploitation of parasocial relationships in a retail setting.
They probably ignore them, as they should - the same problem exists everywhere, from ATMs to door keypads to stores to self-checkout to tapping your card on stuff, etc.
> initially, replace cashier labour costs with new expensive machines,
Labor, like energy, is conserved in the system. It might be easier to counter proliferation of those systems if the narration was focused less about companies replacing labor on their side, and more on the fact that this labor gets transferred to the customers, who are now laboring for free for the company, doing the same things that used to be done better and faster by a dedicated employee.
Today, you need to bypass "do you have the app", "do you want fries with that", "do you want to donate", "are you sure you don't want fries?" and a couple more.
All this is exactly what your parent comment was saying: "To a business, the user interface isn't there to help the user achieve their goals - it's a platform for milking the users as much as possible."
Regarding sanitation, not sure if they are any worse than, say, door handles.
What I wrote earlier, about business seeing the interface as a platform for milking users, applies just as much to human interface. After all, "do you want fries with that?" didn't originate with the Kiosks, but with human cashiers. Human stuff is, too, being programmed by corporate to upsell you shit. They have explicit instructions for it, and regular compliance checks by "mystery shoppers".
Now, the upsell capabilities of human cashier interface are limited by training, compliance and controls, all of which are both expensive and unreliable processes; additionally, customers are able to skip some of the upsells by refusing the offer quickly and angrily enough - trying to force cashiers to upsell anyway breaks too many social and cultural expectations to be feasible. Meanwhile, programming a Kiosk is free on the margin - you get zero-time training (and retraining) and 100% compliance, and the customer has no control. You can say "stop asking me about fries" to a Kiosk all day, and it won't stop.
It's highly likely a voice chat interface will combine the worst of the characteristics above. It's still software like Kiosk, just programmed by prompts, so still free on the margin, compliant, and retrainable on the spot. At the same time, the voice/conversational aspect makes us perceive the system more like a person, making us more susceptible to upsells, while at the same time, denying us agency and control, because it's still a computer and can easily be made to keep asking you about fries, with no way for you to make it shut up.
It will depend on the material of the door handles. In my experience, many of the handles are some kind of metal, and bacteria has a harder time surviving on metal surfaces. Compare that to a screen that requires some pretty hard presses in order to get registered inputs from it, and I think you'd find a considerably higher amount of bacteria sitting there.
Additionally, I try to use to my sleeve in order to open door handles whenever possible.
About as unsanitary as opening the door to get to the kiosks. The kiosks get wiped down more than the door.
LLMs, if used at all, aren't aware enough to even know what the software can do, and many actual chat UIs are worse than that!
My "favourite" design pattern for chat UIs is to invite you type, freely, whatever you like, then immediately enter a wizard "flow" like it's 1991 and entirely discard everything you typed. Pure hostility.
Dynamically generated interactive UIs are something people are barely beginning to experiment with; we don't know if current models can do them reliably for realistic problems, and how effort has to go into setting them up for any particular product. At this point, they're an expensive, conceptual solution, that doesn't scale.
and I'm surprised that people don't bring this visualisation up more often.
And would this company spend billions of dollars for this infinitesimally small increase in convenience? No, of course not; you are not the real customer here. Consider reading between the lines and thinking about what you are sacrificing just for the sake of minor convenience.
Talking to an x-Model (still not AI), just like talking to a human, has never been, is not now, and will never be faster than looking at an information-dense table of data.
x-Models (will never be AI) will eat the world though, long after the dream of talking to a computer to reserve a table has died, because they are so good at flooding social media with bullshit to facilitate the sales of drop-shipped garbage to hundreds of millions of people.
That being said, it is highly likely that is an extremely large group of people who are so braindead that they need a robot to click through TripAdvisor links for them to create a boring, sterile, assembly-line one-day tour of Rome.
Whether or not those people have enough money to be extracted from them to make running such a service profitable remains to be seen.
"I stamp the envelope and mail it in a mailbox in front of the post office, and I go home. And I’ve had a hell of a good time. And I tell you, we are here on Earth to fart around, and don’t let anybody tell you any different...How beautiful it is to get up and go do something."
The Rome trip is even more absurd. Part of the fun of a trip is figuring out what you want to do.
This seems like a product aimed at the delusional, self important, managerial class.
Ok but that does not mean others share the same opinion. Try doing a walk in for a fancy restaurant on the weekend, see how that goes?
Dont be limited with these examples.
How about Airline booking, try different airlines, go to the confirmation screen. then the user can check if everything is allright and if the user wants to finish the booking on the most cheapest one.
What's the goal of technology ? Automate everything so that we don't have to live anymore ? We might as well build matrix pods at that point
I am not looking forward to a trip booked for wrong dates with the hotel name confused/hallucinated for a different one.
But… what if I told you that AI could generate an context-specific user interface on the fly to accomplish the specific task at hand. This way we don't have to deal with the random (and often hostile) user interfaces from random websites but still enjoy the convenience of it. I think this will be the future.
Maybe it could read HN for me and tell me if there is anything I'd find interesting. But then how would I slack off?
Let the bot deal with the ads, the cookie banners, the upsells, "newsletters" and all of the other web BS we deal with.
The bot clicks through the front door of the website, just like us. No APIs, no keys, no nothing.
"Hey Siri, grab me a bottle of slow release 500mg Vitamin C from either Amazon or Walmart, whichever has the best deal. Kthx"
Now with the power of AI we have added back in that middle man to countless more services!
If these are your pain points in life, and they're worth spending $500b to solve, you must live in an insane bubble.
Are these tasks really complex enough for people that they are itching to relegate the remaining scrap of required labor to a machine? I always feel like I'm missing something when companies hold up restaurant reservations (etc.) as use-cases for agents. The marginal gain vs. just going to the site/app feels tiny. (Granted, it could be an important accessibility win for some users.)
I feel like people keep trying to push voice/chat interfaces for things that just flat out suck for voice? The #1 think I look for on a doordash page is a picture of the food. The #1 thing on a stubhub page? The seat map of course. Even things that are less visual like a flight booking, not only is it something that is uncommon and expensive so I don't want to fiddle with some potentially buggy overlay, I can scan a big list of times and numbers like 100X faster than an AI can tediously read them out to me. It only works if I can literally blindly trust that it got the best result result 100%, which is not something I think even a dedicated human assistant could achieve all the time.
I've personally tried using voice to input address in the google nav, and it never understands me, so I've abandoned the whole idea.
I think I sympathize with your feeling but I don't agree with the premise of the question. Do you have or have you ever had a human personal assistant or secretary?
An effective human personal assistant can feel like a gift from God. Suddenly a lot of the things that prevent you from concentrating on what you absolutely must focus on, especially if you have a busy life, are magically sorted out. The person knows what you need and knows when you need it and gets it for you; they understand what you ask for and guess what you forgot to ask for. Things you needed organized become organized while you work after giving minimal instructions. Life just gets so much better!
When I imagine that machines might be able to become good or effective personal assistants for everyone … If this stuff ever works well it will be a huge life upgrade for everyone. Imagine always having someone who can help you, ready to help you. My father would call the secretary pool to send someone to his office. My kids will probably just speak and powerful machines will show up to help.
And I'm not knocking the idea of agents. I can certainly imagine other tasks ("research wedding planners", "organize my tax info", "find the best local doctor", "scrape all the bike accident info in all the towns in my county") where they could be a benefit.
It's the focus on these itty bitty consumer tasks I don't get. Even if I did have a personal assistant, I still can't imagine I'd ask them to make a reservation for me on OpenTable, or find tickets for me on Stubhub. I mean, these apps already kind of function like assistants, don't they?, even without any AI fanciness. All I do is tell them what I want and press a few buttons, and there's a precise interface for doing so that is tailored to the task in each case; the UX has been hyper-optimized over time by market forces to be fast and convenient to me so that they can take my money. Using them is hardly any slower than asking another person to do the task for me.
If the machines are smart enough, shouldn’t they be able to build better interfaces to existing software?
With that aside, it seems like there are two things at play in this demo:
1. Pixel-tuned GPT-4o
2. “Agent” in prod (supervisor loop + operator loop)
Will be interesting to see if they open those up as separate tools in the future, or if they let this fall to the wayside like GPTs, Dalle, etc.
There is no "intelligence" in any of this. Just a whole lot of automation.
Unlike this demo, it uses a simpler interface (Vim bindings over the browser) to make control flow easier without a fine-tuned model (e.g. type “s” instead of click X,Y coords)
I was surprised how well it worked — it even passed the captcha on Amazon!
I understand that in theory it's more flexible, but I always imagined some sort of standard, where apps and services can expose a set of pre-approved actions on the user's behalf. And the user can add/revoke privileges from agents at any point. Kind of like OAuth scopes.
Imagine having "app stores" where you "install" apps like Gmail or Uber or whatever on your agent of choice, define the privileges you wish the agent to have on those apps, and bam, it now has new capabilities. No browser clicks needed. You can configure it at any time. You can audit when it took action on your behalf. You can see exactly how app devs instructed the agent to use it (hell, you can even customize it). And, it's probably much faster, cheaper, and less brittle (since it doesn't need to understand any pixels).
Seems like better UX to me. But probably more difficult to get app developers on board.
If your product has a well documented OpenAPI endpoint (not to be confused with OpenAI), then you're basically done as a developer. Just add that endpoint to the "app store", choose your logo, and add your bank account for $$.
And this kind of seems like an assistant for those.
ChatGPT voice and real-time video is really a beautiful computing experience. Same with Meta Ray Bans AI (if it could level up the real-time).
I'd like just a bulleted list of chats that I can ask it to do stuff and come back to vs watching it click things. E.g.: Setup my Whole Foods cart for the week again please.
Not to be that guy, but where's the evidence for this? People have been telling us that voice interaction is the future for many, many years, and we're in the future now and it's not. When I look around -- comparing today to ten years ago -- I see more people typing and tapping, not fewer, and voice interactions are still relatively rare. Is it all happening in private? Are there any public metrics for this?
That's it. The problem is getting Postmates to agree to give away control of their UI. Giving away their ability to upsell you and push whatever makes them more money. Its never going to happen. Netflix still isn't integrated with Apple TV properly because they don't want to give away that access.
I'm not convinced this is the path forward for computers either though.
With this approach they'll have to contend with the agent running into all the anti-bot measures that sites have implemented to deal with abuse. CAPTCHAs, flagging or blocking datacenter IP addresses, etc.
Maybe deals could be struck to allow agents to be whitelisted, but that assumes the agents won't also be used for abuse. If you could get ChatGPT to spam Reddit[1] then Reddit probably wouldn't cooperate.
[1] https://gizmodo.com/oh-no-this-startup-is-using-ai-agents-to...
I expect many more sites to adopt login requirements. This has the added benefit of more tracking/marketing data.
Of course, already with LLM-powered search we see growing number of people doing the selfish/idiotic thing and blocking or poisoning user-initiated LLM interactions[0]; hopefully LLM tools following the practice above will spread quickly enough to beat this idea out of peoples' heads.
--
[0] - As opposed to LLM company crawlers that scrape the web for training data - blocking those is fine and follows the cultural best practices on the web, which have been holding for decades now. But guess what, LLM crawlers tend to obey robots.txt. The "bots" that don't are usually the ones performing specific query on behalf of users; such bots act as User Agents, neither have nor ever had any obligation to obey robots.txt.
Open API interoperability is the dream but it's clear it will never happen unless it's forced by law.
AI’s are (just) starting to devalue the moat benefits of human-only interfaces. New entrants that preemptively give up on human-only “security” or moats, have a clear new opening at the low end. Especially with development costs dropping. (Specifics of product or service being favorable.)
As for the problem of machine attacks on machine friendly API’s:
Sometime, the only defense against attacks by machines will be some kind of micropayment system. Payments too small to be relevant to anyone getting value, but don’t scale for anyone trying to externalize costs onto their target (what all attacks essentially are).
That convention, implemented well to distribute & decentralize spike impacts, would force any direct overuse attack to take on significant financial risk. While essentially not costing anyone else.
It might still be damaging to availability, but as a service provider I would rather get paid handsomely for periods of being too overwhelmed to service my legitimate customers than not.
The main benefit is having machine interfaces, but those kinds of attacks being heavily disincentivized.
In nearly every case (that an end user cares about), an API will also have a GUI frontend. The GUI is discoverable, able to be authenticated against, definitely exists, and generally usable by the lowest common denominator. Teaching the AI to use this generically, solves the same problem as implementing support for a bunch of APIs without the discoverability and existence problems. In many ways this is horrific compute waste, but it's also a generic MxN solution.
You answered your own question. You have to build the ecosystem if you want to have the facilities your comment outlines.
Whereas the facilities are already in place for "Operator"-like agents.
Even better, it will be difficult for companies who object to users accessing their resources in this fashion to block "Operator"-like agents.
It's certainly a much more difficult approach, but it scales so much better. There's such a long-tail of small websites and apps that people will want to integrate with. There's no way OpenAI is going to negotiate a partnership/integration with <legacy business software X>, let alone internal software at medium to large size corporations. If OpenAI (or Anthropic) can solve the general problem, "do arbitrary work task at computer", the size of the prize is enormous.
When iPhones first came out I had to use Safari all the time. Now almost everything has an app. The long tail is getting shorter.
You can even have several Operator-y apps to choose from! And they can work across different LLMs!
I sincerely hope it's not the future we're heading to (but it might be inevitable, sadly).
If it becomes a popular trend, developers will start making "AI-first" apps that you have to use AI to interact with to get the full functionality. See also: mobile first.
The developer's incentive is to control the experience for a mix of the users' ends and the developer's ends. Functionality being what users want and monetization being what developers want. Devs don't expose APIs for the same reason why hackers want them - it commodifies the service.
An AI-first app only makes sense if the developer controls the AI and is developing the app to sell AI subscriptions. An independent AI company has no incentive to support the dev's monetization and every incentive to subvert it in favor of their own.
(EDIT: This is also why AI agents will "use" mice and keyboards. The agent provider needs the app or service to think they're interacting with the actual human user instead of a bot, or else they'll get blocked.)
For example, by guiding your users to app instead of website, you immediately "lost" 30% of your potential revenue from them. On paper it sounds like something no one would every do. But in reality most developers do that.
OS specific, but Apple has the Scripting Support API [0] and Shortcut API for their app. Works great.
[0]: https://developer.apple.com/documentation/foundation/scripti...
(or, perhaps, agents could use web accessibility tools if they're set up, incentivizing developers to make better use of them)
Instead we need apps that have a human interface for users, and a machine interface for models. I've been building web applets [2] as an lightweight protocol on top of the web to achieve this. It's in early stages, but I'm inviting the first projects to start building with it & accepting contributions.
[1]: https://unternet.co/
Even when it comes to shopping, most of the time I spend is in researching alternatives according to my desired criteria. Operators doesn't help with that. o1 doesn't help because it's not connected to the internet. GPT-4o doesn't help because it struggles to iterate or perform > 1 search at a time.
If the operator is told to find the most nutritious eggs for weight gain, the agent can refer to the nutrient labels (provided by Instacart) and then make a decision.
In the future we might well stumble into those kind of spaces on the net accidentally, look around briefly, then excuse ourselves back to the well-lit spaces meant for real people.
When OpenAI first came out with their first version of GPTs, it was all based on open APIs.
Now they are moving away from it more and more. This means they want to control the market because they don't want to base it on an open standard.
It's such a shame!
As long as people make money from meatspace eyeballs looking at banners, these agents will be actively blocked or restricted just like all other scrapers.
I dream of a world where I can specify annoying things to me and build a perfect search for any house, that understands how I think about money, how I think about my family, and what I love and really extends how I interact with the world.
Instead of hardcoding some automation using Selenium, this would be a great option for automating repetitive tasks with legacy business software, which often lacks modern APIs.
If they demonstrated a big value add like automating CRM a smaller subset of professionals would be absolutely awed but most people would be scratching their heads wondering what it’s good for.
"Create a meme coin for a currently popular meme. Promote it on X and Instagram. Hold onto half the issued coins. When the market cap exceeds US $10 million, start dumping the coins. Send the proceeds to an account in the Bahamas."
Did OpenAI release anything beside this product? Any benchmarks at least to compare?
It feels like OpenAI is betting on the fact that they have a nice UI?!
Paper: https://arxiv.org/abs/2312.08914
Of course, when billions are on the line, and eventually when people are off running their own, ethics get murky. Longer term, I suspect we'll see a global push toward requiring cryptographically asserted human-ness through deep integration with hardware. Cloudflare (and, I'm sure, others) has done some significant work in this domain [1].
[1] https://developers.cloudflare.com/fundamentals/reference/cry...
At the moment it seems kinda useless - to be truly useful it should support querying across multiple sites simultaneously IMO.
For example, the query "Order Joseph Joseph Platform from amazon.com" is something I could easily do faster myself. All the examples shown in their video are similarly simple and don’t showcase much value.
What would be impressive is if you could ask, "Order Joseph Joseph Platform from the cheapest site," and it could compare the total cost (including product price, shipping, and VAT) across all relevant Amazon domains, eu.josephjoseph.com and other shops that ship to my country. Then we’d really be talking.
> While individuals can perform such tasks on their own time at no extra cost, Operator can do so less reliably for US-based ChatGPT Pro subscribers, who pay $200 per month.
Sounds amazing, sign me up :D
E.g.
"Every month, log in to LES.com and pay the current balance. If the balance exceeds $500, alert me before paying."
Ever since GPTs, "Operator" looks quite frankly gimmicky.
We already have a way to build websites for machines: it's called APIs. And frankly, I think that's a better answer for "hooking LLM into website" -- the things which make APIs hard for humans (discomfort, inconvenience, low discoverability, technical complexity) aren't really problems for LLMs.
This kind of thing could just require the devs on one side to maybe clean up a bit of markup (which I even doubt), and the entire universe of potential consumers on the other side.
1. Better bot detection
2. Agents voluntarily providing bot-detection signals (navigator.Bot?)
3. Websites blocking bots entirely.
- double down on influencing the agent’s choices/decision by presenting highest-bidder options (for stores, restaurants etc).
- find a way to figure out the unique user ID behind an agent (which can be trivially done once the agent logs in to say google or facebook etc)
- partnership with AI agent companies to offer their service for free / cheaper in exchange of a complete lack of privacy. For instance the agent could agree to all the cookie policies in the world and let the ad tracker know who you are.
I wish I had something more curious to say.
Other than "Operator/Agent, please surface all sites using prompt injecting and just go ahead and cancel my account, and send a complaint to the appropriate authorities BBB/Reddit/CANSPAM"
For most things I would wish to have done I must login to the site.
If I read the demo right, there is a browser that OpenAI runs that the agent is acting on, not the local browser on my computer. (which could be logged in)
"Agent book me a flight Boston tomorrow before 11am"
[1] Last Question By Isaac Asimov https://users.ece.cmu.edu/~gamvrosi/thelastq.html
"Agents", or something like that.
https://research.google/blog/google-duplex-an-ai-system-for-...
Given how much work is going into "safety", I wonder if this is a field in which less safe open source could overtake the premium models.
While it was scaling, someone(s?) smart went and did a UXR study.
Turned out even if you had a 100% success rate (i.e. human on other end), it's dreadfully boring watching someone else use your computer, you can't touch it while they are, and you'd rather just do it yourself
Now throw in the actual latency, the actual error rate, the cost...I am very comfortable saying this is a waste of time, product-wise.
i want to know how to train for a particular use case for my company. lets say i want to train a model for learning how to use JIRA. how would i go about building this?
Wonder what’s changed recently..
Sounds pretty hollow when some of these companies are quite transgressive, and sometimes revel in it as a form of marketing, and OpenAI has been supplying Israel during an obviously genocidal campaign.
Product also seems like SeleniumBase but much more expensive to execute.
Just keeping sending us money...
But LLMs? Those have already scraped all the data they're going to, and bigger models have less and less impact. They're about as good as they're ever going to be.
There's a problem of "target fixation" about the capabilities and it captures most conversation, when in fact most public focus should be on public policy and ensuring this has the impact that the society wants.
IMO whether things are going to be good or bad depends on having a shared understanding, thinking, discussion and decisions around what's going to happen next.
Let them make their AI if we have to. Let them use it to cure cancer and whatever other disease, but I don't think we should be allowing it to be used for commercial purposes.
Public information and the ability for public to analyze, understand and eventually decide what's best for them is by and large the most relevant aspect. Your decisions are drastically different if you learn soemthing can or cannot be avoided.
You can't dissallow commercial purposes. You can't even realistically enforce property rights for illegal training data, but maybe you can argue that the totality of human knowledge should go towards the benefits of the humans, regardless of who organizes it.
However there's a lot that can be done like understanding the implications of the (close to) zero-sum game that's about to happen and whether they are solvable in the current framework and without a first principles approach.
Ultimately, it's a resource ownership and resource utilization efficiency game. Everyone's resource ownership can't be drastically change but their resource efficiency utilization can as long as the implications are made clear.
They're not even the first movers here: Anthropic's been doing this with Claude for a few months now. They're just the first to combine it with a reasoning-style model, and I'd expect Anthropic to launch a similar model within the next few months if not sooner, especially now that there's been open-source replication of o1-style reasoning with DeepSeek R1 and the various R1-distills on top of Llama and Qwen.
The problem for them is making enough money for the training runs (where it seems like their strategy is to raise money on the hope they achieve some kind of runaway self-improving effect that grants them an effective monopoly on the leading models, combined with regulatory pushes to ban their competitors) — but it seems very unlikely to me that they're losing money serving the models.
So for example, there is a ratio of 10% paid users and 90% free users (just random numbers, not real). If they want more revenue they want to add more paid users, for example double them. But this means that free users needs to double too. And every real free user requires a lot of compute for his queries. Nothing to be cached, because all are different. No way to meaningfully offer "limited" features because the main feature is the LLM, maybe it is previous gen and a little bit cheaper to run, but not much. They can't offer too old software, because competitors will offer better quality and win.
So there is no realistic way to bring costs down. Analysts forecast they actually need to increase prices a lot to meet OAI targets, or it needs to have a financial intravenous line constantly, like the 500B$ announced by Trump.
If you look at their business strategy, it's top notch, anchor pricing on the 200, 20 sweet spot, probably costs them on average $5/mth to server the $20/mth customers, Take your $50m a year marketing budget and use it to buy servers, run a highly optimized "good enough" model that is basically just wikipedia in chatbot and you don't need to spend a dime on marketing if you don't want to, amazing top of funnel to the rest of your product line. I believe Sam when he says they're losing money on the $200/mth product, but it makes the $20/mth product look so good...
They're really playing business very well.
Mid-term, I believe the only real moat is going to be human labor - that is, RLHF and other funny acronyms that boil down to getting people to chat with the model and rate how they feel about its answers.
Software improvements (architecture, training process, inference) are always one public paper or leak away from being available for free to anyone. Hardware improvements will spread too, because NVIDIA et al. would prefer to sell more chips than less chips. Meanwhile, human labor is notoriously expensive, only getting more expensive as economic conditions of people improve, and most importantly, whatever "spark" of human intelligence/consciousness there is, this is where it cannot be automated away - not until we get to human-level AGI.
Human labor is the one thing that you can only scale by throwing more money at it - which is why modern businesses seek to remove it from the equation as much as possible. Hell, the whole pursuit of AGI is in big part motivated by hope of eliminating labor costs entirely. Except, in this one pursuit, until AGI is reached, labor is a critical resource that has no substitute.
That's my mid-term prediction. Long-term, we'll hit AGI and moats won't matter anymore.
I think you are just wrong here and so everything that follows is wishful.
"Plenty" may be vague, but it's not wrong.
"The data is the moat" was a pretty common belief a few years ago, but not anymore.
this was, is and is going to be a constant thing with every AI company
This won't mean humans can't earn wages by selling their labor. But it will mean that human intellectual labor will be not valuable in the labor market. Humans will only earn an income by differentiated activity. Initially that will be manual labor. Once robotics catches up, probably only labor tied to people's personality and humanness.
There seems to be some sort of consensus among legal teams of big American tech companies that the EU is sometimes not worth it for now since OpenAI are not the only one not offering some service in the EU (I'm thinking meta.ai).
Still I haven't been able to find information about what exactly prevents them from selling anything in the EU.
If you're collaborating with the companies the agent is supposed to interact with, why not just have it hook into an API rather than jumping through hoops to interface with their GUI? I don't get it.
Or, HTTP POST.
In 18 month, apps will have APIs for "agentic browsing" ™OoTheNigerian ;)
And you will not need to give anything control over your browser. I you will merely connect your app to OpenAI or any other client.
Same thing with this.
The other day I was writing some code to compute some geometric angles, and I was getting 2 different results for what I though was the same angle, but in fact I didn't realize that these angles should not be equivalent. No LLM was able to tell me the issue, they just said double check my work.
important work. glad to hear they're investing $500B in this space instead of stuff like, I don't know, making the planet livable for our grandkids
The point is that this kind of tool is potentially a real labor-saver for those who are trying to act responsibly within their sphere of influence.
it's a tragedy that with all the ways that humanity could be improved -- and there's a very wide range to choose from -- this is what we choose to spend half a trillion on
Browserbase just launched one of those as a demo
(And yeah, they just got half a trillion).
Edit: Downvote all you want, reality won't change.
Oh, what happened with "Scarlett Johansson will take down OpenAI because she invented speaking like a woman", literally nothing.
What about "AI will never replace Hollywood actors".
What about that time when "OpenAI was done because Ilya was leaving". What a bunch of fools, lmao. I'm not a fan of Sam Altman, but I'm also not deluded.
They haven't got half a trillion either, look up more details. It's a wish they have, funding right now amounts to around $100 billion
It's simple, with trillions at play, if it's so easy to steal OpenAI's game, why has no one done it yet? Don't "argue" about it, just go and grab the money, it's easy, right?
Unless it's misrepresented, this looks like an earmarked VC fund
For instance, the American Recovery and Reinvestment Act (ARRA) of 2009, which allocated approximately $787 billion, was estimated to have created or saved between 2.4 and 3.6 million jobs by early 2011. This translates to a cost of roughly $218,000 to $328,000 per job
In contrast, a study summarized by economist Valerie Ramey in 2011 found that each $35,000 of government spending produced one extra job.
Federal Highway Administration estimated that every $1 billion in federal highway and transit investment supports approximately 13,000 jobs for one year, equating to about $77,000 per job.
https://en.wikipedia.org/wiki/American_Recovery_and_Reinvest... https://www.nber.org/system/files/working_papers/w17787/w177... https://www.fhwa.dot.gov/policy/otps/pubs/impacts/
If the goal is to create data centers for more AI training, you can rest assured that depends on creating as few jobs as possible in order to keep labor costs down and have more to spend on hardware and energy.
- OpenAI hasn't gotten any money yet, not even $100b - OpenAI is releasing this feature after Anthropic - AI has yet to replace any significant fraction Hollywood actors
I'm not disagreeing that these things might happen, but you have to be cognizant of the fact that you're talking projections of the future -- not assessments of the here-and-now.