Are we essentially doomed?
We don't even know how to align models, but even if we did, apparently undoing that alignment if trivial.
Really I'm looking for any argument that lays out a scenario where this works out.
We don't even know how to align models, but even if we did, apparently undoing that alignment if trivial.
Really I'm looking for any argument that lays out a scenario where this works out.