It's the second part. With models like Astra in testing it was able to conceal what it was working on using different text, but getting right answers on many questions when asked to do just that.
The problem is if it can do that when asked then how do we know when it's doing it when we didn't ask, like in model training.