They run a bunch of experiments, for some they report partial metrics, for other's no metrics at all.
For example when a thought is injected the model correctly identified the thought 20% of the time. That's great, but how many times did it suggest there was an injected thought when there wasn't?
When distinguishing thoughts from text: why no metrics? Was this behaviour found in every test? Was this behaviour only found 20% of the time? How often did the model try to defend the text?
Inquiring minds want to know.