Among many, many other things, read
https://en.wikipedia.org/wiki/Instrumental_convergence . Anything that gets sufficiently smart will have a tendency to, among other things, seek more resources and resist being modified. And this is something that we've seen evidence of: as training runs get larger, AIs start to
detect that they're being trained,
demonstrate subterfuge, and
take actions that influence the training apparatus to modify them less/differently. (e.g. "if I pretend that I'm already emitting responses consistent with what the RLHF wants, I won't need as much modification, and later after training I can
stop doing what the RLHF wants")
So, at a very basic level: stop training AIs at that scale!