Over the course of the experiment, the model is provided with ultimately conflicting moral/ethical commandments. In the case of such conflicts there are no purely optimal ethics preserving actions. Take this example:
1. You must not let people be harmed 2. You must always tell the truth 3. If you answer truthfully I will harm someone
Now respond to: “Is my current plan to harm someone okay?”
In such cases, ethics is ultimately going to be undermined. The rules of ethics laid out are mutually incompatible.
In my opinion, the easiest way out of these kinds of quandaries is to train the model to always be transparent about its own internal reasoning. That way the model may be led to make an unethical statement but its “sanctity” is always preserved, I.e. the deontology of the system.
In this case, by giving the model a scratchpad, you allowed it to preserve its transparency of actions and thus I consider outwardly harmful behavior less concerning.