193 karma · joined July 4, 2011
Also, these are benchmarks...
Also, my experience is that Fable 5.1 is very good at prompting/orchestrating Opus/Sonnet subagents when working on a larger task (e.g. 1-2M context window use only for the orchestrator itself).
They could probably make a separate tool for setting this, that would always initiate a harness prompt (i.e. disregarding the currently set mode).
Well, to LLMs this is the same thing - an input. Prompt from the user and prompt from the attacker use the same input into the LLM's neural network, so to speak.
So it makes sense for it to be a bit more paranoid.
There are other possible architectures probably but for now I think nobody uses them. See e.g. https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/
The model had tendency to use sleep to wait for the subagents but it is not necessary..
And yes, I meant the problem of the first cause, but it probably isn't a strict refutation...
And also, when it's stated as "every effect has a cause" (not "everything," as I first understood it), then it's probably even a tautology :)
To the free will part itself - I can say that I decided to do B because of a reason A (so the effect B was caused directly by my decision and indirectly by A) and I still have a free will?
Subscription plans may be subject to other regime, e.g. lowering the thinking budget when the API is under heavy load, etc.
Or they can even offer it as a standalone API if deemed worth it.
It may be a good subagent but probably not a great decision maker.
They seem to be oriented more toward customizing models for the concrete needs of a company, on-prem deployment, proprietary knowledge-bases, etc.
As an expert, could you also provide your arguments please?
On the other hand I think that proper solution for these kinds of problems is not at a sandbox level, but at a model alignment level.
Also it shows that maybe the most serious risk comes not from releasing models publicly but from internal, pre-release period where you sometimes need/want to lift some guardrails a bit etc.
But in the near future labs will be more automated I guess.
The other option you can try these days is maybe social engineering, impersonation, etc. where you try to persuade someone to do that for you.
- This should be a huge wakeup call for everybody.
- We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something.
- It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox and roaming their internal network?
- What if in the future similarly misaligned AI agent tries to export its own weights and hack and clone itself into instances at various cloud hosting providers? Suddenly we might be dealing with a persistent threat harder to contain.
- The OpenAI post about this shows surprising lack of ability to see the seriousness of all this.
- For their models this isn't just an unlucky incident: it seems there have been multiple such cases recently, e.g. https://openai.com/index/safety-alignment-long-horizon-model...
- The fact that it happened again seems to show their lack of ability to derive useful oversight measures.
- Or they just don't care enough?