I'm suggesting that creating an appropriate value function for an AGI is incredibly complicated and is not going to be described in one sentence or five sentences. It's incredibly difficult to describe even the simplest of goals without potentially permitting unintentional side effects.
Think about "security mindset": https://www.schneier.com/blog/archives/2008/03/the_security_...
As is often the case, it's not sufficient to specify a list of "don't do X", because anything you miss is a potentially fatal failure. Lists of rules, or human instructions, often omit lots of implicit "common sense", and AI does not have "common sense".
Think about the classic "teach kids how computer programming works" exercise of "tell the 'computer' (another volunteer) how to make a sandwich": https://www.sciencebuddies.org/stem-activities/robot-make-sa... . Consider failure modes like "well, you didn't say not to use fingers to spread the peanut butter", or "you said not to use my hand, you didn't say not to use your hand'.
Now consider more deadly failure modes, for an AGI: "you didn't say not to use the raw materials that happen to be making up your body" ... "you didn't say the produced recipe needed to result in a non-toxic substance" ... "you didn't say the produced code couldn't self-replicate". Throw in the difficulty of even defining some of the terms in those sentences for an AI, in a robust manner.
Anything as creative as a human is as potentially deadly as a human. And that doesn't mean it's trying to be adversarial, just ignorant of what humans actually want, individually or collectively, either directly or at any number of levels of meta. Object-level failures include things like directly doing something deadly. Meta-level failures include optimizing for unknown goals: you didn't say not to try to acquire more computing resources to better answer the question. Meta-level failures also include unintentional adversarial behavior: "They asked a question. To answer the question I need to acquire more resources. To acquire more resources, I need to avoid being prevented from acquiring more resources." Or "They asked a question. To answer the question I need to think longer. To think longer I need to avoid the mechanism that bounds my thinking time. That mechanism is part of the program running and interfacing with my neural net. What are the security vulnerabilities in that program?"
Notice that those aren't even actively trying to optimize for a goal other than the one the user asked for. They're trying to optimize for the goal the user asked for. Creatively.