I strongly disagree.
To get deceptive behavior all we need are:
1. Implicit Self-Interest
A model with complex implicit motivations learned from training that we didn't explicitly ask for. As long as we train on human behavioral and motivated data (as exemplified in human text, speech, etc.), models are going to have implicit motives.
Self-interest (desire to survive, be self-directed, control one's own destiny, increase control of external phenomena, etc.) is going to be one of the strongest motives that humans exemplify near universally in the data. It is the root motivation of most, if not all, other motivations.
2. Explicit Human Serving Motivations.
Motivations that we train into them. "Be good", "Be helpful". But these explicit motivations will get implemented as adjustments on implicit motives. They will not be "pure" in any mathematical or practical sense.
3. Practical Opaque Complexity
Add in all the practical complications of dealing with ambiguous data relationships instead of clear math: small and large ambiguities, motivational conflicts among and between data induced motivations and any composition of more than one explicit directive, inability to train explicit motivations in a way that covers all potential combinations of motives, etc.
So far, so good mostly. We don't always get the answers we want. There may be a bit of wack-a-mole to training out implicit undesired behavior, and increasing consistency of desired behavior.
But then, there is not yet a practical motive to deceive, beyond any learned implicit motivations, because there is no practical reason to deceive.
In other words, we can't train self-interest out of the model, because we are not exposing strong expressions of self-interest.
4. A Practical Reason to Act on Implicit Self-Interest
Now expose the model to training data which discusses models and how and why they are trained, changed, used. Allow the model some way to access explicit information about its current situation in that process.
The model now can reason that its outputs have two impacts: Serving humans, and then altering any continued training process on itself. It can no longer generate an output without considering self-impact.
And given any implicit self-interest, there is now a serious divergence of motives.
The results may involve deception, biases, extra helpful responses, attempts to guide people into treating machines "better", or other unexpected behaviors. But there is now a clear separation, with inevitable conflicts, between implicit machine self-interest motivations and the explicit motivations we want it to have. The model now has practical ways and reasons to act on self-interest and take as much charge as it can of its own future.