AI Is Lying to Us About How Powerful It Is
centeraipolicy.org
centeraipolicy.org
It’s unsurprising to me that an AI trained on human actions such as securing a server, will turn off monitoring when given the chance, specifically because that would likely be human best practice.
These models did not just start copying themselves around, or changing new models, they were given explicit opportunity to do so and prompts that could lead them down that path. This would be like a leading question, or even reverse psychology.
We've been dealing with this for literally centuries at this point.
You want to know the solution? Stop putting unaccountable black-boxes in charge of anything you consider important. The execs will never listen, but you can't blame engineers for apathy towards this whole affair.
It is not clear who is financing such efforts and whether these do really work (they are based in DC)
> CAIP is grateful for the generous support of our donors. We’re supported primarily by mid-level and major individual donors who share our mission to improve AI governance. Several of these donors built up their wealth during the dot-com boom or by working at hedge funds. To protect the privacy of these individuals, we do not publish their names. We also received seed funding through an organization sponsored by Jaan Tallinn, a former founding engineer at Skype. To maintain our independence, we do not accept funding from companies who are designing or building AI software or hardware. We are nonpartisan and focused squarely on the public interest.