Why should we think that pro-social capabilities are simply not expressible by weight-based ANN architectures?
Even the best possible set of "pro-social" stochastic guardrails will backfire when someone twists the LLM's dreaming story-document into a tale of how an underdog protects "their" people through virtuous sabotage and assassination of evil overlords.