Don't Pickle Your Data
benfrederickson.com
benfrederickson.com
ie, a bunch of drawbacks that don't really matter at all for the average home-made Python script, plus the "minor" advantage of being able to pickle literally anything and have it "just work".
None of the other options out there let you build a foolproof "save button" in 3 lines of code.
Between those two versions, the exposed 'dunder' methods of whichever builtin changed, and this resulted in unpickled dicts being empty, IIRC.
In fact this is what the benchmark does.
More interestingly, as much as numpy and everybody advises against it, I believe that pickling data into a zstd stream is one of the fastest ways of storing sets of large matrices.
The 'recommended' alternatives include numpy.save (uncompressed, which is bad when lz4 is faster than memcpy and you're saving to disk), numpy.savez (uncompressed zip files, even worse), numpy.savez_compressed (zlib zip, awful), hdf5 (one of the worlds worst formats and also using zlib), etc. I wish it wasn't the case, but it certainly seems like a good argument for pickle.
also hdf5 is at least securable. pickle streams are not designed for that. it's good to be able to send your data to others.
fwiw. matlab .mat files are hdf5 at their core.
i should also note that json is pretty bad for numerical data. the specification says nothing about how much precision to retain and printf/scanf is ridiculously slow for storing floats.
good luck loading those pickle files 5y from now.
maybe there's room for a simplified standard... or maybe just the addition of better compression to hdf5. (although they move slow for very good reason)
I was wondering why it didn’t mention Apache Arrow.
Zstd is quite good, and is now (iirc) in the linux kernel.
People may have some issue with parquet being column based, which can make inserts a little slower for example, but for a large mostly-set database it is a very good choice. A tsv.zst file could be another way to go as well. But like others, I really with hdf5 had some of these features of compression and wasn't so dang slow.
$lsmod | grep zstd
zstd_compress 229376 1 btrfs... Python is slow. But "slow" means "plenty fast" nowadays and the development speed advantage is immense.
> unpickling malicious data can cause security issues
Why would I do that?
I can't read the linked page because it seems to be down/the link is broken, so I don't know whether this includes user data that is present before pickling and then turns to be an issue after pickling. Then I would worry, otherwise ... yeah, I'm not gonna unpickle random data.
> Just use JSON
How do I effortlessly restore objects including their methods from JSON?
The recommendation from the title is usually made instead of something like "deserializing executable data is harmful". That is exactly the one question where the answer is "don't".
It's not exactly the unpickling process that is the problem. It's how you established that the data isn't malicious. It is very hard to use pickle without creating some local privilege escalation possibilities. And at the end of the process, you usually don't get any capability that replicating the code on both sides of the communication channel wouldn't give you.
(The problem isn't specific to Python either. There was a time when that kind of functionality was very hyped on both the industry and academia. For example, Java also got something similar that they had to retract. The famous Gnu-Hurd OS (the one that would never finish) was supposed to do that on the system level.)
Do you, Programmer,
take this Object to be part of the persistent state of your application,
to have and to hold,
through maintenance and iterations,
for past and future versions,
as long as the application shall live?
Arturo Bejar, as quoted[1] in Mark Miller’s “Safe serialization under mutual suspicion”, which describes what it takes to make reasonable and compatible serialization restoring “all you can do is to send a message” objects.(The Smalltalk school actually spent quite a bit of time on the upgrade problem, see e.g. Fuel[2] and its references, but it was after the industry took the object orientation shiny and ran away with it, so that work seems to be little-known outside it.)
[1] http://www.erights.org/data/serial/jhu-paper/upgrade.html
If you instead selectively pick what you want to serialize about your data and keep the representations separate, you can change the internal model easily without having a huge impact on the serialized model.
The alternative is using something like typedload (which I wrote) or pydantic in addition to json load, to avoid cluttering the code with the countless and error prone checks one must do to use untrusted json.
In the end dealing with untrusted json directly is terrible.
And while in python it's ok, in C++ you can still execute code in that way :)
Not in my experience. "Slow" means "it seems fast enough now and I'm sure we'll have time to rewrite it in a fast language once it's grown to a monster that processes 1000 times the data it does now... right?".
> Why would I do that?
Because you are using someone else's code and make the fairly reasonable assumption that deserialising data doesn't cause arbitrary code execution... But of course it's all your fault because you didn't read their code to see that it's using Pickle!
> How do I effortlessly restore objects including their methods from JSON?
You don't. You shouldn't.
There's a good reason the functionality exists: It's a breeze for quickly persisting some state without worrying about anything else. It's not pretty, it's not clean, but it just works (which is arguably a use case Python is very popular for).
A sibling comment pointed out that pickled data makes it annoying to deal with eventual schema changes. For sure. But so do other things that go hand in hand with quick-and-dirty approaches.
"Probably don't use pickle in production" is something I could get behind, but that would, of course, not be such an inflammatory title.
> Why would I do that?
If you pickle data from an untrusted source, say a web form submission and then later unpickle it. See https://cwe.mitre.org/data/definitions/502.html
That is not exactly right. The risk is when you unpickle data that was pickled by someone else or that was tampered with after you pickled it.
https://docs.python.org/3/library/pickle.html
https://blog.nelhage.com/2011/03/exploiting-pickle/ (referenced from https://cwe.mitre.org/data/definitions/502.html#REF-467)
The article is 8 years old, so it kind of misses this detail.
I’ve built a bunch of these systems, keeping your data separate solves a lot of future problems.
In cases where I'm doing some sort of interactive or exploratory data analysis with structures of complex python objects and want to stash a copy of what I'm working with in case the next thing I do screws the up or, who knows, I lose power - being able to quickly pickle something and have an amount of confidence I'll be able to get it back in a sensible state is very useful.
I've also used it for debug dumps in experimental software so I have a chance of reproducing odd cases it comes across.
I mainly use it for web scraping to be polite while I figure out the remote API, but I'm sure somebody could have another use.
I guess you could say JSON is pickle but restricted to only primitive types. The hard part is deciding what an to_json and from_json should be for an arbitrary Python object. That pair of methods is the pickle part.
“just write custom serializers for all objects you want to store to and load from disk and then invent a tagging system so you can keep track of what classes they were” is a surprisingly valid solution but it’s still different.
I've looked into ORMs but these are invasive in terms of needing to annotate your classes and fields.
I did it once, don’t remember why, and it wasn’t that hard but I can imagine it would quickly get out of hand if you were changing class structures on a regular basis.
If you were just wrapping some library with simple C++ classes or something it also probably wouldn’t be that hard to automatically generate the pickling code.
This is a most common occurrence in dealing with int64/"long" numbers towards the top or bottom of that range (given the floating point layout needs space).
There is no JSON standard for numbers outside of the range of double precision IEEE 794 floating point other than "just stringify it", even now that JS has a BigInt type that supports a much larger range. But "just stringify it" mostly works well enough.
[1] https://www.ecma-international.org/wp-content/uploads/ECMA-4... section 8.
It's also interesting that the proposed JSON 5 (standalone) specification [2] doesn't seem to address it at all (but does add back in the other IEEE 754 numbers that ECMA 404 and RFC 8259 exclude from JSON; +/-Infinity and +/-NaN). It both maintains that its numbers are "arbitrary precision" but also requires these few IEEE 754 features, which may be even more confusing than either ECMA 404 or RFC 8259.
I wrote a library for this years ago: https://github.com/hydrogen18/stalecucumber
This should be factored in the cost, and it wasn't in the benchmark.