Make your schema permissible and your code paranoid, it will pay off later. Build a data linter if necessary, but don't tie the schema.
Make your schema permissible and your code paranoid, it will pay off later. Build a data linter if necessary, but don't tie the schema.
I phrase this as: If it doesn't make sense to do math on it, it's not a number. What does adding one to a customer account number mean? Absolutely nothing -- you get a completely different account number. So it's not a number, but a numeric string.
Namely, when constructing a protobuf, theoretically, there might be two different ways: (A) first gather all the fields, and then construct the protobuf from these fields; (B) first construct an empty protobuf, and fill in the fields as necessary. The actual protobuf uses (B) - which is convenient in most cases, because when you start constructing a protobuf usually you don't have all the data ready yet.
However, with required fields, this means when you construct the protobuf it starts with all required fields missing - i.e., an invalid state!
I'm not sure what's the best way to fix it, because it would be infeasible to rewrite all the code to gather all the fields and then construct the protobuf - also it will be hugely inefficient in many cases. However, I feel the "no required fields" rule is essentially a null pointer (the "billion dollar mistake") in disguise - the actual problem is that the API doesn't enforce type safety.
Imagine you have an innocent `required` field. You have a producer and a consumer of that field that communicate over the wire. (or instead of the wire, imagine a database).
You send or store an instance of that protobuf. Now let's say that you want to make the field optional (or remove it). With an already-optional field, this is easy. You stop setting it, and maybe eventually you clean it up.
With a required field, however, you can't do that. If any of your clients don't have the newest schema version, you can't unset the field (so imagine that you support mobile clients who may never update). Or if there's middleware you don't know about that introspects your proto. Even if you do the dance right and update your server and client before not setting the new field, you could crash outdated middleware that you didn't know about. Whoops!
Or with the database, you now need to dual write or something complex because if you need to roll-back to an older version, you'd be unable to read the protos that don't include the required field.
Required doesn't do well over time. It has nothing to do with setting the values.
Changing required to optional isn't a magic fix for protocol compatibility. If it were (for your limited use case) you can just make that change to the protobuf client side as it doesn't affect the wire representation/interpretation.
Right, there's the rub. `required` means that anyone who deserializes your proto falls into this category. That's a much larger group than "anyone who reads a specific field". So the list of clients now is forced to include any and all middleware that may read your proto (imagine a routing layer or some kind of analytics system or whatnot).
(Note also that there's lots of ways to make reading a field that is empty fallback to doing some reasonable non-catastrophic behavior, required doesn't let you do those things).
Making invalid states unrepresentable is basically the process of taking human-checked invariants and turning them into type-checked invariants. This reduces the likelihood of bugs and guides humans to use the system correctly.
Required now means requires forever because people can't migrate safely. But technically you can change a protocol descriptor from required to optional, which is invalid (usually, in a distributed non-transactional system (the common kin) but nothing stops you from doing it. So why not make required forever? Well, do you really want to commit to anything forever?
Required usually cannot be deprecated.
Because to be sure of that you can remove it, you need to be sure that every storage system and every piece of middleware and every since thing that links your proto anywhere in the world that you might care about is upgraded, otherwise if they encounter a new message they'll crash.
If you only have a single client and server, and you control both, this is doable. If you don't have that though, you cannot.
For example, in the case of the message bus they say "And even though the message bus doesn’t care about message content", and later on "The right answer is for applications to do validation as-needed in application-level code." Strict schema and validation is most helpful for application developers, not some middleware routing code. Was it not possible for them to write a parser that doesn't fully validate the message for use-cases like this?