GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction.
Would love to hear your feedback!
GPT-Live-1 is the first version of a new generation of models, and we believe the full-duplex architecture + delegation enables entirely new ways of human-AI interaction.
Would love to hear your feedback!
Looking at the 30,000-foot view of how society is set up: laws, economic system, employee incentives, etc, do you suppose it matters what the individual contributors think? I say this not to absolve anyone of responsibility, but to point out the obvious outcomes of our incentives across the strata (polity -> shareholders -> boards -> C-suite -> employees)
I will bet you dollars to donuts, somewhere inside OpenAI is a frequently-used revenue dashboard, but not for loneliness - if anything, OpenAI will make horny models and tout itself as a solution to loneliness, a la character.ai - if that earns them more money.
I think assuming people will use it as a tutor/learning tool is.. way too optimistic. A small fraction will, but the majority will just view things like a second language as something not worth learning.
I think there's a logic leap here that ignores most of what big companies know about how to gey people addicted to things
Did something bad? Better ablate yourself of the responsibility of holding people to account for making it _worse_! Acting this way just makes it seem more like you regret the blood on your hands because it has dirtied your shirt, and not because you’ve done something actually bad, otherwise you’d have at least some degree of guilt and reticence to see things get worse.
The morality of an action isn’t based on actual consequences because the future isn’t known in advanced. All we can do is act on the perceived consequences of our actions, and if we think those are good, pursue them.
The loneliness epidemic is driven by companies maximizing keeping their customers engaged with their screens, something OpenAI is wont to do. Knowing that the company wants customers engaged and that this will do that, and also knowing that that plays into the loneliness epidemic by substituting human interaction, makes it far different than getting married and then maybe or maybe not getting divorced.
If you're talking about ChatGPT in general (as I guess you are) I think the Jury's still out on whether this will have a net positive or negative on society.
Right now, it's leaning towards negative, but there are optimistic futures to be had at the same time.
As someone who likes to think that they would turn down a lot of tech jobs due to moral dilemmas - for me, currently working on AI wouldn't be one.
For me at-least it's a tool to learn quicker, and reduce friction on projects I otherwise may not have dived into. As with all new technology, we're still in this grace period / lack the bigger picture we now have on other, now known to be destructive technologies.
Awesome. Are you guys able to share anything about the model architecture? I've been interested lately in split-transformer RVQ-based conversational agents, e.g. via stuff like https://arxiv.org/abs/2412.10208 (ResGen) and https://arxiv.org/abs/2603.18090 (MOSS-ITT) and of course Moshi (https://arxiv.org/pdf/2410.00037).
Intuitively, decoupling semantic and audio-timeslice-space generations with coupled but distinct histories is right model architecture, not just for these sorts of assistants, but for domains like robotics too.
One big gap I've run into for UX is most realtime voice harnesses wait for a full response from tools, and at most support the model filling the dead air until then
It'd be a game-changer to be able to have the model start replying with partial information streamed from the tool call, then seamlessly continue with additional information.
Some people have standards in what they build.
You can't continue a conversation if it hinges on information that takes inherently takes 30+ seconds to generate. Previously even models that "handle it" can only play for time.
You can try and simulate this by messing with model context mid-conversation but that breaks down reliability massively as the model loses track of what it's talking about.
1. The voice model delegates to one agent.
2. The voice model delegates to multiple agents, and keeps track of tasks.
3. The voice model delegates to an orchestrator agent, which then delegates to sub-agents and keeps track of tasks.
YMMV depending on the exact product experience you care about, because there is a tradeoff between latency and layers of delegation.
Our current implementation is backed by one model, but you can imagine this getting much better with time.
And because the voice is so frictionless to talk to, I asked about what company owns the building, then that company's industry, then how that industry works in this particular country etc. I probably wouldn't have bothered going down a rabbit hole like this if I'd had to type. Voice is much easier than typing.
Anyhow it's fun! Thanks for making it!
I'm currently on the 20 $/mo subscription and using codex meaningfully, and i'm loving this.
I am considering bumping my subscription to the 100 $/month and this might be the reason i switch, BUT: i really envision me using this also through other means as well (eg: agents like openclaw/hermes) in agentic ways.
Will this be supported?
I can make OpenAI stuff the center of my agentic AI life, but I need it to be interoperable.
i.e. how will full duplex & delegation enable/enhance desktop flows w/o corresponding leaps in UI.
One group's expectation of interruption for pleasant conversational flow can be just as off-putting as another's expectation of patient silence.
Are we seeing any conversational layer integrated with codex soon?
Here's a voice I really enjoy listening to. https://elevenlabs.io/app/voice-library?search=zNsotODqUhvbJ...
If I could have Christopher Lee, or Stephen Fry, I would.
- The videos felt scripted and dishonest