Protobuffers Are Wrong (2018)
reasonablypolymorphic.com
reasonablypolymorphic.com
The truth is, a lot of these critiques are valid. You could redesign protobuf from the ground up and improve it greatly. However, I think this article is bad. Firstly, it repeatedly implies that protobuf was designed by amateurs because of its design pitfalls, ignoring the fact that it's also clear that protobuf grew organically into what it is today rather than deliberately. Secondly, I do not feel it is giving enough credit to the idea that protobuf simply was not designed to be elegant in the first place. It's not a piece of modern art. It is not a Ferrari. It is a Ford F150.
What protobuf does have right is the stuff that's boring; it has a ton of language bindings that are mostly pretty efficient and complete. I'm not going to say it's all perfect, because it's not. But protobuf has an ecosystem that, say, Cap'n Proto does not. This matters more than the fact that protobuf is inelegant. Protobuf may be inelegant, but it does in fact get the job done for many of the use cases.
In my opinion, protobuf is actually good enough. It's not that ugly. Most programming languages you would use it from are more inelegant and incongruent than it is. I don't feel like modelling data or APIs in it is overly limited by this. Etc, etc. That said, if you want to make the most of protobuf, you'll want to make use of more than just protoc, and that's the benefit of having a rich ecosystem: there are plenty of good tools and libraries built around it. I'll admit that outside of Google, it's a bit harder to get going, but my experiences on side projects have made me feel that protobuf outside of Google is still a great idea. It is, in my opinion, significantly less of a pain in the ass than using OpenAPI based solutions or GraphQL, and much more flexible than language-specific tools like tRPC.
For a selection of what the ecosystem has to offer, I think this list is a good start.
https://github.com/grpc-ecosystem/awesome-grpc#protocol-buff...
I don't even feel this way. I can complain for days about most technical things, including things Google has made (e.g. Golang), but not protobufs. They're just good. They don't totally replace JSON of course, but for the intended use case of internal distributed systems, they're solid.
Then use thrift, which does plain http. But most of the design decisions in the article are the same there as well. Turns out there's really only one way to do this.
But tech at Google tends to be pragmatic in the specific way that protobuf is. It's not perfect, it doesn't fit neatly into a grand ideology, but it is (at least within Google itself) simple enough, easy enough to understand, portable and fit for purpose. In a similar way to bazel, it's full of components worthy of criticism, but those fixes get made when they become the ecosystem's most pressing issues, and not before.
I've known this forever, but when my company started doing outreach to customers, the point was really driven home.
Find a problem. Solve it. That is it. If the solution is popular enough it will pick up momentum. There will always be naysayers. It might not be the best solution, or even the right one, but it IS a solution, which is better than what others are doing.
Think something is bad? Fix it. Submit a PR if it is open source. Don't complain for the sake of complaining unless you are willing to show how it should be done.
Note that I'm not one of those open source jerk holes, however I do recognize that unless I can show I can do better, my opinions are meaningless.
Edit: kentonv's rebuttal says this too. "This article appears to be written by a programming language design theorist who, unfortunately, does not understand (or, perhaps, does not value) practical software engineering." And when my angry hat is on, this is also how I would generalize programming language design theorists.
It's not a pro level system designed by a serious telecom team.
ASN.1 is more optimized in other ways that make sense, sacrificing some flexibility in favor of performance because the ends are expected to be more coordinated, but that's separate.
https://news.ycombinator.com/item?id=21871514 (211 comments)
https://news.ycombinator.com/item?id=18188519 (298 comments)
In particular, Kenton's rebuttal at https://news.ycombinator.com/item?id=18190005 is worth a read.
Once you start doing rpc, the demand goes so up. Cap'n'Proto's ability to return & process future values is such an exciting twist, that melds with where we are with local procedure calls: they return promise objects that we can pass all around while work goes on.
Kenton popped up a couple months ago to mention maybe finding time for mutli-party support, where I can for ex request say the address of a building & send the future result to another 3rd party system. Now this is less just a way to serialize some stuff, & more a way to imagine connecting interesting systems.
Thanks!
edit: seems, not possible as of now - https://github.com/capnproto/capnproto/issues/478
> Cap’n Proto gets a perfect score because there is no encoding/decoding step.
How do they achieve this?
No, there isn't. Not any more than there is a decode step anyway.
The truth is that encoding, decoding, and validation all occur when you call the accessor methods for specific fields. If you never actually call the getter for some field, it won't be decoded nor validated at all.
If you are going to be reading the entire message tree, then zero-copy is only an incremental improvement on something like Protobuf -- the use of fixed-width values may make decoding faster (or slower, in some cases). The real magic is when you want to read just one field of a 10GB file. Just mmap() the whole thing and treat it like a byte array, and it'll be efficient. Cap'n Proto will not attempt to scan the whole file, only the parts that you explicitly query.
(I'm the author of Cap'n Proto.)
Also if your language's bindings don't preserve the relationship between x and has_x and don't play nice with recursion schemes, use a better binding generator! You don't have to use the default one.
Protocol buffers were written by Jeff Dean and Sanjay Ghemawat. Whether you see issues with the implementation or not and whether they are fit for your use case or not are valid discussions, but if your argument starts with "omg these idiots couldn't design a simple product" then I'm already reading the rest of it with a huge grain of salt.
Reading through the rest of the post though, it is pretty clear that it is a case of trying to use what is a standard for data transfer over the network to describe your entire application's object model. It isn't the right tool for the job, and doesn't have to be.
If it is being forced on you then, well, that's a complaint for your management, not the technology itself.
True, but a lack of self awareness is never a good look.
Making all fields in a message required makes messages into product types but loses compatibility with older or newer versions of protocol. Auto-creating objects on read dramatically increases chances that a field removal in future version of protocol will be handled properly. Auto-creating on write simplifies writing complex setters, etc...
You could argue that those problems could be solved better (and I might agree, protobufs are one of my least favorite serialization protocols) but not even acknowledging reasons and pretending the designers didn't know better makes writer seem either ignorant or arguing in bad faith.
If the protobuf schema code generator had to translate these suggestions into efficient wire representations and efficient language representations, it would be more complex than half the compilers of the languages it targets.
I do mourn the days when a project could bang out a great purpose-built binary serialization format and actually use it. But half the people I hire today, and everyone in the team down the hallway that needs to use our API, can no longer do that. I'm lucky if they know how two's complement works
And really that's how the network stack on Linux and Co. is implemented (modulo taking care of endianness)
Till one day someone wakes up and wants a UI in C#. Which is also not a big deal, but you have to either hand roll the serialization code or build a header file parser and generate it. And then someone needs to add a field - so we just tack it onto the end of the struct. This works fine as long as you have the length represented out of band somehow. If not then you get to make struct FooBar2 with your extra field. Then a year later someone has the great idea that it would be great to send text, and now you're making a variable length packet or just always sending around structs with a bunch of empty space and a length field (which is also not too bad). But wait, now the UI team totally needs that data in JavaScript for their new Web UI, so you're back to either generating or hand jamming dozens of structs. All the while tamping out bugs in various places where someone is running old code against new data or old data against new code - all of which have various expectations. Or that microcontroller that is blowing up because it just doesn't like that unaligned integer that is packed into the struct.
Not that I'm bitter about that life - it just involved a lot more troubleshooting protocols over the years than I'd have liked. Anyway, those are problems that Protocol Buffers helps solve. But as long as you're using just C and not changing much - packed structs are quite lovely.
I don't think people should think too much about "what if we change language", etc. because (1) that's unlikely to happen, (2) you have years ahead of you, (3) it may be simpler to convert structs (not the most complicated thing in the world amd supported in most languages in one form or another) into whatever else when/if you actually need it than to overengineer now 'just in case'.
And versioning only works when someone remembers to increment the version.
With packed structs you don’t get the reverse compatibility, so schemas and endpoints have to be immutable which is not the end of the world but it can make changes more difficult.
Yes, exactly this. I don't understand the vitriol here. Protobufs work fine for a wide variety of purposes and where the criticisms matter people use a different tool.
> I do mourn the days when a project could bang out a great purpose-built binary serialization format and actually use it. But half the people I hire today, and everyone in the team down the hallway that needs to use our API, can no longer do that. I'm lucky if they know how two's complement works
Would be great to work at a place where everyone knew how the machine worked, but the vast majority of developers entering the workplace since about 2000 have learned Java and web sh*t exclusively.
Protobuf seems to encourage this… instead of a hard boundary where your serialization logic ends and your business logic begins, every project I’ve worked on that uses protobuf tends to blur the lines all over the place, even going as far as to make protoc-generated types the core data model used by the whole code base.
Numbers are all rough, but seemed to be in the same ballpark in all the sources I checked (like https://thescalers.com/development-deep-dive-how-many-softwa...).
You can't use numbers and raw logic because it has various levels of inaccuracies and things you haven't accounted for especially when the numbers and data come from times prior to the layoffs.
Reports from recruiters and people who are unemployed are telling me things that are significantly different. There is a fundamental change in the job market.
I do agree with some of the schema modelling criticisms in this article, but the ultimate thing to understand is: protocol buffers were invented to allow google to upgrade servers of different services (ads and search) asychronously and still be able to pass messages between them that can be (at least partly) decoded. They were then adopted for wide-range data modelling and gained a number of needed features, but also evolved fairly poorly.
IIRC you still can't make a 4GB protocol buffer because the Java implementation required signed integers (experts correct me if I'm wrong) and wouldn't change. This is a problem when working with large serialized DL models.
I was talking to Rob Pike over coffee one morning and he said everythign could be built with just nested sequences key/value pairs and no schema (IE, placing all the burden for decoding a message semantically on the client) and I don't think he's completely wrong.
I think required fields are fine, provided that you understand that "required" means "required forever". If you're already using protos this isn't exactly a brand new concept. When you use a field number that field number is assigned forever, you can't reuse it once that field is deprecated. It requires a bit more thoughtful design, and obviously not all fields should be required, but it has value in some places.
Quite a statement from someone that spent less than a year at Google.
Personally I really appreciate protobuf; it has saved me and teams I've worked with a ton of time and effort.
Putting this article on my ignore list.
Microsoft has DCOM and other .net remoting mechanisms, which may even work over the Internet.
Next/Apple has Distributed Objects: https://developer.apple.com/library/archive/documentation/Co...
And databases have had it since forever, you throw in some SQL and you get back some data over defined protocols like TDS. You can even invoke some small programs via stored procs.
I have no idea WTF SOAP is though …