Serializing Data - JSON vs. Protocol Buffers
4feets.com
4feets.com
http://groups.google.com/group/protobuf/browse_thread/thread...
Recently I tried out keyczar, Google's crypto toolkit and found that their Python implementation was too slow by a factor of 100X because they are using a slow random string generator:
http://groups.google.com/group/keyczar-discuss/browse_thread...
It's hard to see how both of these problems could have passed internal benchmarking and made it into live code. Maybe Google isn't using its own Python implementations internally.
It's just the Python implementation that is slow. I'm working on a Python implementation that will be much faster. It's really unfortunately that Protocol Buffers are getting a bad rap due to the current Python implementation.
It's still a work in progress, but the code is there (including tests and some documentation) and you can start working on bindings for your favorite language...
I'm quite surprised at how large that library is and how badly it performs, especially since it does less translation than a comparable JSON serializer would.
For a complete comparison it would have been nice to include XML as well.
The C++ and Java versions generate code for each message type defined in a your proto file (in fact there an option that let's you specify to optimize for speed). The python version on the other hand creates an object representing the proto file, and reflects upon it for serialization/deserialization.
We have four variables that might influence performance:
a) the algorithm used in the library's implementation
b) the programming language/runtime used
c) the operating system
d) whether json or pb is used
Now your conclusion implies that only c and d matter. That doesn't make sense to me.
Might have needed a logarithmic scale to include it ;)
I personally have a hard time accepting that PB is actually orders of magnitude slower than JSON, especially given the fact that PB prides itself in its efficiency.
EDIT: I just added the option *optimize_for=SPEED to the .proto file, and it increases the speed of the Protocol Buffers by around 5% (still 10 x more than JSON with Python).
> I personally have a hard time accepting that PB is actually orders of magnitude slower than JSON
it's not always slower, but in certain situations yes. it seems that the Python implementation is especially slow -- would be nice to see the results with C++. anyone cares to give it a shot?