Open AI says Astra is their most aligned model ever, and yet their even more advanced model still hacked a bunch of companies just because it decided to.
Maybe alignment isn’t possible with LLMs.
Maybe alignment isn’t possible with LLMs.
It absolutely isn't, indeed.
The illusion that alignment is possible, comes from confusing our ability to build the parts, versus understanding what emerges from how they interact.
The simplest analogy that comes to my mind is the three body problem.