> Google's implementations, at least C++ and Java, are a bunch of bloated crap (or maybe they're very good, but for a use case that I haven't yet encountered).
As someone who has been working on protobuf-related things for >10 years, including creating a size-focused implementation (https://github.com/protocolbuffers/upb), and has been working on the protobuf team for >5 years, I have a few thoughts on this (thoughts are my own, and I don't speak for anybody else).
I think it is true that protobuf C++ could be a lot more lean than it currently is (I can't speak to Java as I don't work on it directly). That's why I created upb to begin with. But there's also a bit more to this story.
The protobuf core runtime is split into two parts, "lite" and "full". If you don't need reflection for your protos, it's better to use "lite" by using "option optimize_for = LITE_RUNTIME" in your .proto file (https://developers.google.com/protocol-buffers/docs/proto#op...). That will cut out a huge amount of code size from your binary. On the downside, you won't get functionality that requires reflection, such as text format, JSON, and DebugString().
Even the lite runtime can get "lighter" if you compile your binary to statically link the runtime and strip unused symbols with -ffunction-sections/-fdata-sections/--gc-sections flags. Some parts of the lite runtime are only needed in unusual situations, like ExtensionSet which is only used if your protos use proto2 extensions (https://developers.google.com/protocol-buffers/docs/proto#ex...). If you avoid these cases, the lite runtime is quite light.
However, there is also the issue of the generated code size. Unlike the runtime, this is not a fixed cost, but is proportional to the number of messages you use. If you have a lot of messages it can quickly dwarf the size of the runtime. For this reason, C++ also supports "option optimize_for = CODE_SIZE" which uses reflection-based algorithms for all parsing/serialization/etc instead of using generated code. This means you pay the fixed size hit from linking in the full runtime, but the generated code size is much smaller. On the downside, "optimize_for = CODE_SIZE" has a severe ~10x speed penalty for parsing and serialization.
I have long had the goal of making https://github.com/protocolbuffers/upb competitive with protobuf C++ in speed while achieving much smaller code size. With the benefit of 10 years of hindsight and many wrong turns, upb is beginning to meet and even surpass these goals. It is an order of magnitude smaller than protobuf C++, both in the core runtime and the generated code, and after some recent experiments it is beginning to significantly surpass it in speed also (I want to publish these results soon, but the code was merged in this PR: https://github.com/protocolbuffers/upb/pull/310).
upb has downsides that prevent it from being fully "user ready" yet: the API is still not 100% stable, there is no C++ API for the generated code yet (and C APIs for protobuf are relatively verbose and painful), it has a bunch of legacy APIs sitting around that I am just on the verge of being able to finally delete, and it doesn't support proto2 extensions yet. On the upside, upb is 100% conformant on every other protobuf feature, it supports reflection, JSON, and text format, but also lets you omit these if you don't want to pay the code size.
I hope 2021 is a year when I'll be able to publish more about these results, and when upb will be a more viable choice for users who want a smaller protobuf implementation.