[edit] I'm genuinely asking. I'm guessing you mean in the generated APIs, i.e. the code "weight", but maybe you meant something else?
[edit] I'm genuinely asking. I'm guessing you mean in the generated APIs, i.e. the code "weight", but maybe you meant something else?
Things may have changed since, but AFAIK the C++ implementation would always allocate on the heap for nested messages, and perhaps even for optional scalars. This may be optimal for larger documents, but not for smallish messages (my use case was market data and trading instructions). I measured certain small messages, where an encode/decode pair would take over a microsecond with Google's implementation, but about 50 ns with a simpler one (versus 15 ns for a memcpy).
For Java my experience is mostly with the API itself, which felt very heavy.
Edit: I think a lot depends on your use case. I use protobuf mostly as a 'trusted' protocol. If someone didn't set a required field, I don't care. Some bloat may have to do with verifications that I've never needed.
This is no longer the case if you use arenas: https://developers.google.com/protocol-buffers/docs/referenc...
> and perhaps even for optional scalars
This has never been the case, except for string fields where std::string forces us to allocate.
Ideally we will eventually use std::string_view for string accessors instead of std::string, so that even string data can be allocated on an arena instead of the heap.
I'm was quite surprised you didnt offer your own stringview implementation (or something similar) the last time I looked at protobuf. I'd naively assume that inside Google this could be quite a low-efford high-reward optimization.
We sort of do actually: https://github.com/protocolbuffers/protobuf/blob/master/src/...
The internal version of protobuf lets you switch individual string fields to string_view using [ctype=STRING_PIECE], but migrating the default away from std::string is mainly just an enormous migration challenge.
Internally we also do something slightly nuts: we break the encapsulation of std::string so that we can point it to arena-allocated memory (we then "steal" the memory back before the destructor runs). We can only afford to do this internally, where the implementation of std::string is known. The real long-term solution is to move to string_view.
Arenas have been around for a long while now though...
FWIW if you reuse the same message object for multiple parsings, it will re-use the sub-objects as well, thus amortizing away the allocation cost. Parsing the same message into the same object twice should do zero allocations on the second parse. This is the intended way to use Protobuf for small-size messages.
Apparently the C++ implementation has also grown support for arena allocation more recently. (After my time, so I don't know much about it.)
Take a look at an implementation like Prost, for Rust. It's very similar to what I did (10 years ago by now). Everything is just inline, except when messages can be recursive (which should be rare for most protocols).
The big problem is bloat in memory usage if you parse many differently-shaped messages, requiring the app to implement hacks like only reusing a particular object a certain number of times.
Many messages are have lots of optional sub-message fields, and set only a few of them in any given message. These messages would be huge if everything is inline (especially if the same thing happens in those sub-messages).
I agree that inlining all sub-messages works great for dense schemas, but it assumes too much about the schema to be a good design for a general-purpose proto library I think. Also maps and repeated fields can never be inline.
So it's great that there are different implementations for different use cases. It helps that the wire format is simple and well-documented.
In the case of protobufs, there is no real answer. Protobufs is one of the few network libraries that doesn't operate with zero-copy. That steps over a tangible line in the sand.
No zero-copy for networking? Forced internal heap allocations with only this arena feature after a decade? Sorry no. Protobufs isn't useful for serious network applications.
That's a bit harsh. Protobufs deliver smaller wire size than any of the newer "zero-copy" formats. And many receivers of zero-copy formats will... copy the data into some internal representation. If your protobuf implementation delivers classes that are good enough to work with internally (store in maps, forward, etc) then you don't really lose something; instead you gain, due to no manual conversion layer.
That's not my position at all. In my other comment (https://news.ycombinator.com/item?id=25586447) I explain how I've spent 10 years trying to improve on protobuf C++ precisely because I agree that some of these limitations are unnecessary.
> No zero-copy for networking? Forced internal heap allocations with only this arena feature after a decade? Sorry no. Protobufs isn't useful for serious network applications.
I suppose it depends what you are comparing it to. Almost every JSON library has the same limitations you mentioned, and yet many people find JSON useful for network applications. But I agree that giving users full control over allocations makes a library useful in many more situations.
I think arena allocation is a pretty reasonable solution to the problem. You can use whatever memory you want for the arena (stack, heap, static buffer) and you can constrain it so that no heap allocations are allowed.
Unfortunately protobuf C++ can't live fully within this arena model while it uses std::string for accessors. Hopefully this can be fixed at some point.
JSON is a human-readable format which is hugely advantageous to develop and operate in many settings. Protobufs doesn't have that advantage. Yet we're paying all of the same costs to structure the data with both. That's enough to posit JSON as a net winner over protobufs.
> Unfortunately protobuf C++ can't live fully within this arena model while it uses std::string for accessors.
Forcing std::string as a container for core components of a networking API (specifically for an accessor) is an exemplary demonstration of a lack of seriousness in a library.
im not super experienced with c++, so please forgive the maybe obvious question, but why is that?
(im guessing it has to do with allowing control over where allocations come from?)
Would you go through all of the effort to rig an application with userspace networking only to return its results inside an `std::string` -- one of the few std containers where one can't even control the allocator if they wanted to? That's absurd, if anything.
Google does all of that when it makes sense, and yet uses it to push protobufs. That’s the entire point: you are calling it absurd and non-serious, but this only shifts my opinion on you, not on Google.
But you seem to be arguing that Protobuf is bad because it is not well suited to certain use cases, dismissing all other use cases as "non-serious". That is offensive.
In fact, I can just do that now: Protobufs is bad. You seem to agree, because you haven't even advocated for it -- just flatbuffers and Cap'n Proto and the phonebook of alternatives. Listen, I don't care who you've worked for or what you've done. As a scientist and engineer I'm totally ready and able to entertain critiques of my prior works without being offended. I'm offended that this is the quality of your participation here.
In C, dealing with memory allocation is such a pain that C programmers still tend to avoid it. But C++, especially post-C++11, makes dynamic memory allocation much, much easier to deal with. (And GC'd languages, obviously, are easier still.) There is still a performance cost, obviously, but that cost almost never matters in application programming use cases, and is even negligible in many systems programming cases.
I do think Protobuf does too much allocation, but saying that libraries should not allocate memory at all is an outdated view.
The good news is that flatbuffers[1] is a reasonable replacement for most of my use cases. In particular being able to mmap() them directly is a wonderful thing that you can't do with protobufs in addition to being very allocation sparse.
No, modern phones are certainly not constrained in the way I meant, and I don't think you could call them "embedded" either. The common programming languages used on phones are very memory-allocation-friendly.
> There's few things I hate to see more than a flat profile from memory allocation or cache misses.
I think you may be arguing a different point, or a different level of extremity of the point. Reducing memory allocation to optimize performance is a fine thing that everyone does. The person I was replying to, though, seemed to be asserting that libraries should completely avoid allocating memory for themselves.
> The good news is that flatbuffers[1] is a reasonable replacement for most of my use cases. In particular being able to mmap() them directly is a wonderful thing that you can't do with protobufs in addition to being very allocation sparse.
Yeah... I'm the author of Cap'n Proto, which has the same property, and predates Flatbuffers.
I would hardly classify Dalvik or ART as "allocation friendly", they don't perform escape analysis and if you do it constantly you'll be in a world of constant hard GC pauses. Multiple times over the years I've had to build free-lists in Java to avoid this specific problem.
Same for C++ if you use one of the built-in generic memory allocators. The fastest new/delete are the ones that you don't call.
For what it's worth I tend to agree with the grand-parent thread. The lack of awareness of allocations, cache-invalidation via indirection are a significant contributor to why we see software clawing back hardware wins across the years on these platforms.
I don't know what you're doing, but I've generally found free-lists to be a net performance negative in Java libraries. Time and again, I've been called in to "optimize" Java code that uses them, and usually by simply removing them I can get rid of the performance problems entirely.
Flatbuffers in particular excels here since once you have the ByteBuffer in memory you can immediately start accessing data without needing to do any extensive parsing.
Even the official Android docs are very explicit[1] on the allocation point. Allocations are not cheap and even with generational collection you still will blow the 16.6ms frame window for 60 FPS if any of your operations allocate excessively.
[1] https://developer.android.com/training/custom-views/optimizi...
Be that as it may, lots of people still manage to use Android devices.
Allow me to unpack this, because the object interface that all of these serialization schemes present is what increases complexity. It makes sense that the simple implementation would tend toward dynamic memory and not further compound such complexity on their interface. This is a flawed architectural premise taken by these library implementations.
On the receiving side, there is little reason to ever deviate from zero-copy/zero-allocate. These libraries are far too pro-active in deserializing incoming data into native object structures. It's not necessary to present the user with this; they should pique through the results lazily -- query into the data at their own discretion. This is important because it bounds the minimum requisite cost of receiving messages at ... zero. It costs nothing to discard trash. Whereas with these proactive serialization layers, every single incoming attack becomes an exercise against the deserializer. In a zero-copy/zero-allocate -- let's even say, zero-parse mere lexing of the data, software has more flexibility. In practice, that means passing pointers around to parts of messages straight off the wire to application functionality further up the stack.
On the sending side, there is little reason to ever force whole native object representations to conduct a serialization. In other words, don't force the user to build an std::map first and serialize it later. All input is streaming input. Properties do not have to be pre-buffered, they can be streamed. One can model any network serialization this way, sans perhaps canonical representations (sorted JSON keys, etc), and far more efficiently than with requiring arbitrary native structures.
This is not really a rebuttal about the merits for or against dynamic memory in and of itself. I know the research, thread-aware allocators can be pretty good -- even darn good, and the state of the art in GC is nothing to shake a stick at. The problem is that with better design it's just not necessary, and in the end it does have a cost that one should want to eliminate if possible. I'm certain it's quite usually possible, at least more than I see in a list like https://en.wikipedia.org/wiki/Comparison_of_data-serializati... etc
But it does have drawbacks. The encoded size is necessarily a bit larger to allow random access traversal of the raw bytes (though it compresses well). The API to manipulate structures in-place is a little awkward, particularly on the writing side. And, you can't really use the generated types as mutable in-memory state, as people commonly like to do with Protobuf types -- as a result, a common feature request for Cap'n Proto is to support generating "native" structs with the ability to convert between those and the zero-copy types as desired.
Everything is full of trade-offs. I don't disagree with your design preferences but I do object to the extreme line you are taking on them. You said: "Protobufs isn't useful for serious network applications." That is plainly contradicted by the existence of a trillion-dollar company built on said technology.
Sometimes a native object structure is what you want.
Cough.
While with embedded systems there definitely is a big thing about dynamic memory allocation, much as I don't like it, it's not like very popular and successful libraries don't do this. It's a pretty common and accepted practice, and there are standard idioms for how it is done.