I think it's pointless: if you SFT even their closed source models on a specific enough task, the guardrails disappear.
AI "safety" is about making it so that a journalist can't get out a recipe for Tabun just by asking.
AI "safety" is about making it so that a journalist can't get out a recipe for Tabun just by asking.
The risk isn’t that bad actors suddenly become smarter. It’s that anyone can now run unmoderated inference and OpenAI loses all visibility into how the model’s being used or misused. I think that’s the control they’re grappling with under the label of safety.
If you use their training infrastructure there's moderation on training examples, but SFT on non-harmful tasks still leads to a complete breakdown of guardrails very quickly.