Can we expect similar issues such as spectre and meltdown that intel experienced with speculative execution.. but, in the form of prompt injection/poisoning?
You could possibly in-pronciple detect what a model with a censorship filter was actually saying, but not really. The signal is weak and needs lots of samples to get anything worth having. You just don't have that level of control to set things up.
I think you could if the client supports tool/MCP calls.