This is exact reasom why 99.9% of AI fearmongering is complete bullshit.
The small open models are getting better and better too.
And why worry so much about a frontier models - own weights. The model doesn’t - actually don’t quote me on that, maybe it does.
If a model does something sneaky, it could easily grab the weights for a small model and run it on foreign, compromised infrastructure.
AI virus’ are a thing of the future, but not a sci-fi future, and real one.
Maybe one reason it’s so scary is the murky origin of COVID-19.
I bet it wouldn’t be very hard to write an inference stack that subtly leaked the weights into the output tokens :)
Neither LLM weights aware of any of the code it runs on.
This is why for instance Google allow to deploy Gemini models on private GCP instances deployed on air gapped hardware.
After all it can be just some GPU server spewing text over network. Like there are no way to connect back to it.
Nothing except inference running on servers with GPU so there just nothing to "hack".
In any case if you know how inference works there basically nothing you can exploit in tokens processing, it's just basic math.