You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.
You can not tell me that OpenAI doesn’t have the resources or ability to run tests like this in a physically-non-networked environment w/ sufficient compute for its needs.
They don’t even need to be fully airgapped from each other (and is not what I’m suggesting).
But there should be no physical (physical layer; wireless counts) to the internet.
How much of the internet do you have to simulate to know if the model knows it's in training?
Regardless, they (reportedly) _attempted_ to prevent internet access. They just didn’t in a way which can be escaped via software.
Yes, side channel exploits exist in airgapped environments to. But if a model found a way to escape an airgapped environment via non-networked side channel attacks then the correct answer is frankly “shut it down immediately and then thermite any machine it touched”
Anyways, we'll give it to you for only $2T. We need at least that amount to get as far away from here as humanly possible.
Why should it be physically airgapped? Clients won't be doing that.
Is it safe to release such software if it has only been tested in environments where certain major risk areas do not exist?
Appendix C Illustrative safeguards, controls, and efficacy assessments has specific examples like:
- Agent actions are all logged in an uneditable database, and asynchronous monitoring routines review those actions for evidence of harm
- Limiting internet access and other tool access
- Limiting credentials
- Limiting access to system resources or filesystem (e.g., sandboxing)
- Limiting persistence or state
Did we learn nothing from all of those Star Trek holodeck jailbreaks?
And if you are too afraid to test it without guardrails, that probably means it shouldn’t be released.