OpenAIs models since 5.5 were troublesome in ways even a layman like me could reproduce, their testing environments (“sandbox”) downright a showcase of what not to do and they, despite being one of the biggest labs, didn’t observe what any of their models output for weeks after multiple prior incidents. They had multiple warnings, they took not a single precaution.
You operate machinery or software, you are responsible to monitor it.