*I work at OAI.
*I work at OAI.
That's what I've been heads down, HUNGRY, working on, looking for investors and founding engineers pst: https://heymanniceidea.com (disclaimer: I am not associated with heymanniceidea.com)
What plan are you on? I'm starting to wonder if they're dynamically adjusting reasoning based on plan or something.
Opus 4.6 worker agents never asked for permission to continue, and when heartbeat was sent to orchestrator, it just knew what to do (checked on subagents etc). Now it just says that it waits for me to confirm something.
There are bugs, it doesn't work perfectly, but that's just part of testing and refinement at this point.
My initial prompt was just: "let's work on converting this java game to c++ using panda3d. you're a panda3d c++ expert. you will be the agent that owns the project, creating the plan, and the delegating each step to sub-agents that create each system in the correct order."
it created like 17 different tasks and sub agents and opus 4.7 orchestrated it. I did personally validate which rendering engine would be good for the project etc first.
Like I will get Opus to make me an app but it will stop in between because I need to setup the db and plug in the API keys and Opus really can't do that on its own yet
The goal is none. The current situation: everything that matters requires human intervention.
I think the end situation will be that LLMs will be able to perform decently well in a highly controlled and predictable environment.
Why this constraint? A common sentiment I see online (sorry, to group you in) is "[tool] will be capable, actually, but only in a context that trivializes its usefulness."
I think modern post-training like RLVR + inference-time output token scaling can _probably_ scale so the agents can solve any computable task, even when placed in noisy or misconfigured environments. But it won't be economical for a long while. But it already seems largely capable of that today.