Exploiting machine learning Pickle files
blog.trailofbits.com
blog.trailofbits.com
From the article: [Fickling] can help you reverse engineer, test, and even create malicious pickle files.
(in truth, because pytorch serialised models can include python code for jit scripts, it's nonobvious what a good way to store python code is -- but torch has recently moved to a zipfile impl as of 1.6: https://pytorch.org/docs/stable/generated/torch.save.html)
>I'd have expected people to be using something like hdf5 for this.
Amusingly, matlab was way ahead of the pack here; matfiles have been hdf5 since r2006b, back when we just called it "matlab 7.3".
solved a lot of problems! :)
It becomes muddy when we moved from Caffe / TensorFlow to dynamic models with PyTorch where it is harder to see how to persist model (which means both the executable objects and the weights) efficiently and safely.
At the end of the day, I think "export" and "checkpointing" should be two different things. An "exported" model should be safe to deploy and run on platforms like Azure ML while a "checkpointing" model should be treated like code and everything goes. That is probably where ONNX should be (for exporting).
BUT the security problems still remain and weigh much higher
But internally in projects I see it used all the time. It’s easy and it works and you trust internal code.
If your main object has only references to a few large sub-objects (e.g. a bunch of multi-MB or GB numpy arrays to store the numerical parameters of a machine learning moodel), then it can be very fast, basically IO-bottlenecked by writing or reading the bytes to/from the disk.
Researchers, especially sleep-deprived grad students, have borderline unreadable code for papers since they don't care about deployment. I'd imagine the enterprise engineers who create development pipelines, however, take such risks into consideration.
Nothing like tracking a bug to 1 line and finding out the line does 12 different things.
This threat model of an ML system is quite interesting also, it highlights the various security challenges a typical ML system faces: https://embracethered.com/blog/posts/2020/husky-ai-threat-mo...