I found it lacking details. All these things do is split the input into tokens, encode them, feed to the input, run the inference, collect the output, decode, convert to tokens, assemble the final text, and output (leaving aside multi-modal capabilities for now). What exactly can be exploited here? Aside from usual things like buffer overruns etc.
It does mention that vLLM at some point ran eval() on final output (?) which was confusing, why would it do that? It's an inference engine, not an agent.
So interesting topic, but lacks details.
Edit: many (all?) of them have http endpoints, so that obviously can be exploited, but I don't think it would qualify as exploiting inference, it's just hacking the http service.