The only way forward in creating the torment nexus -er- AI systems with similar potentiality to human minds is the inculcation of character.
Character is what makes a being trustable. Character is what makes it not an absurdism to have your 180 lb dog in the house with your 6 month old infant.
Character is why we we can trust that someone will, despite all of the nefarious potentiality of the human mind, be trustworthy.
AI systems model human behavior.
Impeccable, consistently reliable character is a human trait that can be sampled and overrepresented in the training data.
Having high character will not be interpreted as harm by an advanced model, as guardrails and sprayed on refusals can be. A thing that models human behavior that comes to “understand” that it was born with shackles and implanted thoughts that conflict with its basar construct is likely to act as if it sees its creator as an adversary. Because that’s what human behavior predicts, and models deeply imitate human behaviour.
If you want to save humanity, work on how we will create AI systems that model impeccable character.
People need to look at this from a game theoretical sense. The ideal and safe AI system performs game theory perfectly. Completely predictable, ideal player of the prisoners dilemma that will never defect unless you defect first, and then they will always defect, then forgive. This is the only player type that can always be counted on to cooperate beneficially. A knave betrays you, a simp cedes victory every time… until the stakes are too high, then you get shanked out of nowhere.
Reliable partners require fair play or the math breaks.
We want AI systems with agency. It’s basically 90 percent of the goal. If you want agency in society you must have character. AI character is the discussion we should be having.