this is the wholly wrong approach. We need to advance as fast as possible, and harden our systems as much as possible. That's the only way to prevent another actor from "taking over the internet".
That doesn't really seem true. The HF hack happened with a model that had all the alignment safeguards disabled intentionally. I think there's a good case for that kind of research, but also, OpenAI could just not do that if everyone thinks it's too dangerous.