Add extra stuff to a “standard” encoding? Sure, why not
rachelbythebay.com
rachelbythebay.com
> WHY? Oh, why?!
Uh oh. Is this my HN moment?
This is exactly how I implemented it at my company. We had to write many protobuf messages to one file in bulk (in parallel). I did a fair amount of research before designing this and didn’t find any standard for separating protobuf messages (in fact, found that there explicitly isn’t a standard in that protobuf doesn’t care). So I thought rather than using some “special” control character, like a null byte, which would inevitably be not-so-special and collide with somebody else’s (like Schema Registry’s “magic byte”), I’d use something meaningful like the number of bytes the following record is.
As for why I chose varint instead of just picking an interger size, well for one I got nerd-sniped by varint encoding and thought it would be cool to try and implement it in Scala. Secondly, I thought if I chose a fixed size integer, no matter what size I pick, my users will always surprise me and exceed it at least once, and when that happens, kaboom! I wanted to future proof this without wasting 64 goddamn bytes in front of each message, and also I got nerd-sniped, OK?!?
Someone on my team recently shared one of these files outside the company and so I really hope she’s not talking about me but that’s a crazy coincidence if not!
https://protobuf.dev/reference/java/api-docs/com/google/prot...
The fact that protobufs are not self-delimiting is an endless source of frustration, but I know of 2 standards for doing this:
- SerializeDelimited* is part of the protobuf library: https://github.com/protocolbuffers/protobuf/blob/main/src/go...
- Riegeli is "a file format for storing a sequence of string records, typically serialized protocol buffers. It supports dense compression, fast decoding, seeking, detection and optional skipping of data corruption, filtering of proto message fields for even faster decoding, and parallel encoding": https://github.com/google/riegeli
81 A2 70 62 C6 <32 bit be length> <protobuf data>0x8n A map containing "n" key/value pairs -> 0x81 introduces a single key/value pair
b101XXXXX : A string where XXXXX is the string length -> 0xA2 introudces a 2 byte string
0xC6 followed by a 32-bit value N introduces an array of bytes of size N
0x70 0x62 "pb"
So it roughly corresponds to
{"pb": <protobuf data> }/s
64 bytes would be a 512-bit integer. That seems like excessive future proofing for the length of any message that would be transmitted before the Sun runs out of fuel.
Not my thought originally, I heard it from somewhere else but can't find it. Possibly from Foone Turing.
[1] https://www.wolframalpha.com/input?i=log2%28+volume+of+unive...
All digital representations rely on discrete states in hardware, and there’s a finite number of those in the observable universe, so there should a finite maximum number for computers.
We don’t know if there is only a finite number of discrete states in the observable universe. The Planck length is not a discretization.
I'll have to think about this. I want a very physical meaning of the term using values rather than references. One where the largest number that can be expressed using up/down fingers with two hands is 1024, not a sign language reference to a googolplex.
https://en.m.wikipedia.org/wiki/Hyperoperation
I wonder if it would be helpful to restrict the question a bit, maybe something like: what is the largest number for which the number, and also all smaller magnitude integers, can be expressed.
"Measuring the intelligence of an idealized mechanical knowing agent" https://philpapers.org/archive/ALEMTI-2.pdf
"Intuitive Ordinal Notations" https://github.com/semitrivial/IONs
Just as a quick example of why it's a bit absurd - I name that number you just defined $zeta$. Now I make $zeta'$ = zeta^zeta. Or whatever manipulation you like. Adding constraints is addressed in the link.
The GP question was not about encoding, and thus is not subject to compression. The largest number we can measure of anything is a pretty well defined concept.
Plus rather simple things like pi could create rather a long message.
You can, of course, simply read field by field, but few libraries expose the ability to do that simply. And the naive operation then becomes quite problematic indeed.
https://clickhouse.com/docs/en/sql-reference/formats#protobu...
Using a delimiter to mark end of packet. For text protocols we usually use "\n" to be the delimiter. This usually requires some escaping in the packet to make it unambiguous. Two standard protocols for this are SLIP and HDLC.
Length encoding - like what you did.
The downside of the delimiter approach is that it changes the length of the packet - when you are escaping one byte becomes two and that's sometimes a pain if you're doing it in place in memory (less of a problem if you're streaming byte by byte). The big advantage is that it allows for resynchronisation - if you lose a single byte from your stream, or your length byte ever gets corrupted then you're permanently out of sync - the receiver will never again know where the start or end of a packet is. With the delimiter approach, you just lose one packet. So if you're ever doing this for a UART or network stream or something, always do the delimiter approach!
I'm confused. I mean for UART, sure. But network streams is usually sent over a protocol that recovers lost data. Am I reckless for sending length-prefixed data chunks over TCP?
I actually use SLIP for packetisation over TCP anyway, because then syncing for any other logging or diagnostic thing that joins halfway through is easy. Basically I find it to just be a more robust system.
Whoa, that's an acronym I haven't heard in a long time.
SLIP/TCP seems redundant. Your serial encoding IP then sending it over TCP/IP? Is this a tunnel?
You're right that it was originally for segmenting a serial link to send separate IP packets (so you could theoretically have SLIP over TCP over IP over SLIP), but it works just as well in other contexts.
(The TCP checksum is very weak at only 16 bits, but there's usually another layer like TLS that gives you integrity for free so nobody cares)
Shove protobuf into Something Else that does packet delimitation for you. I'm fond of SQLite for offline cases as a richer alternative to sstable.
I shudder to think how well shit must be going for this to merit a Rachel post.
In general, I find adding versions to things allows for much more graceful future redesigns and that is, IMO, invaluable if you're concerned about longevity and are not confident in your ability to perfectly design something for the indefinite future.
The way protobufs handles versioning (old software ignores unknown fields) is far superior and realistically everyone uses 64 bit fixed length sizes everywhere
The idea is so simple that any change would be a misfeature.
The magic means it’s possible to identify the file type. Maybe you’ll add a tool later that operates on multiple types of files.
The version means you can evolve the contents in non-backwards-compatible ways, while maintaining the ability to read/parse the old version.
It’s a pain in the ass to add a magic or version number later; there’s a reason why nearly every file format on the planet has both.
XML
JSON
YML
TOML
<?xml version="1.1" encoding="UTF-8" ?>
The others are serialization formats, like protobuf — they’re not file formats.> they're not file formats
Brb, gonna delete all my files without file formats
If you’re going to use text-based serialization formats as your justification for the decision, however, I’d suggest you look into all the fun bugs, security issues, and weird edges cases that arise from parsers having to make a best guess at character encoding and file format when all you have to work with is the file extension, maybe a byte order mark, and heuristics over the file contents.
Because that's all you would need.
The only issue is if you were to ship a “protobuf” library that emits/consumes your (very much not protobuf) framing format.
Also, a 64 bit frame length would only be 8 bytes, not 64 :-)
Protobuf encoders/decoders commonly implement two formats: delimited format, and non-delimited format. Typically non-delimited is the default, but delimited format is supported by many implementations including Google's main first-party ones. In fact, the Java implementation shipped with this support when Protobuf was first released 15 years ago (I wrote it). C++ originally left it as an exercise for the application (it's not too hard to implement), but I eventually added helper functions due to demand.
Both formats can be described as "standard", at least to the extent that anything in Protobuf is a standard.
So clearly the bug here is that one person was writing the delimited format and the other person was reading the non-delimited format. Maybe the confusion was the result of insufficient documentation but certainly not from a library author doing something crazy.
Merely using the delimited format without any other sort of framing is almost always a bad idea because of precisely the ambiguity TFA discusses.
I'm pretty sure delimited streams are rarely used in the wild instead of something more robust/elaborate such as recordio, which specifically are almost always prefixed with a few magic bytes to mitigate this problem.
Edit: Also, why is there no publicly available recordio specification? Infuriating.
They are somewhat common inside Google at least.
But if the project in question was indeed protobuf.js (see loeg's comments), it clearly distinguishes encode/decode vs. encodeDelimited/decodeDelimited. So I believe the project should not be blamed, and the better question would be why so many people chose to add this exact helper. Well, because Google itself also had the same helper [3] [4]! So at this point protobuf should just standardize this simple framing format [5], instead of claiming that protobuf has no obligation to define one.
[1] https://github.com/protocolbuffers/protobuf/blob/main/docs/t...
[2] https://github.com/tafia/quick-protobuf/issues/130
[3] https://protobuf.dev/reference/java/api-docs/com/google/prot...
[4] https://github.com/protocolbuffers/protobuf/blob/main/src/go...
[5] Use an explicitly different name though, so that the meaning of "encoding/decoding protobuf messages" doesn't change.
Definitely seems to be a routine addition to the standard supported by many libraries.
It only suggests the length prefix and doesn't define the exact encoding at all.
And since some do it in the same way, sometimes it works. The two typical approaches I found in the wild: use varint as a length and do nothing at all. Typically, the second one implies that the user, if they want to send a sequence of messages need to get creative and invent some form of connecting them together.
GRPC is the first kind. So, all those using GRPC rather than straight-up Protobuf are shielded from this problem.
For protobuffs in particular, I have no idea. If you look at the encoding [0], you will see that the notion of submessages are explicitly supported. However, submessages are preceeded by a length field, which makes the lack of a length field at the start of the top-level message a rather glaring omission. The best arguement I can see is that submessages use a tag-length-value scheme instead of length-value-tag. This is because in general protobufs use a tag-value scheme, and certain types have the begining of the value be a length field. This means that to have a consistent and composable format, you would need to message length to start at the second byte of the message. Still, that would probably be good enough for 90% of the instances where people want to apply a length header.
message Foo {
repeated string bar = 1;
}
Any repetition of `09 03 41 42 43` (a value "ABC" for the field #1), including an empty sequence, is a valid protobuf message. In the other words there is no explicit encoding for "this is a message"! Submessages have to be delimited because otherwise they wouldn't be distinguishable from the parent message.Hmm don't they both allow that? Or am I misunderstanding what you mean here?
I guess the interesting (though only occasionally useful) thing about protobuf is if you concatenate two serialized messages of the same type and then parse the result, each repeated field in the first message will be concatenated with the same field in the second message.
The field bytes dont really encode tag and type, they encode tag and size (fixed 32bit, 64bit or variable length)
Protobuf is a TLV format. In that regard, it's not unique at all.
This is misinterpreting what actually happens. "Message" in Protobuf lingo means "a composite part". Everything that's not an integer (or boolean or enum, which are also integers) is a message. Lists and maps are messages and so are strings. The format is designed not to embed the length of the message in the message itself, but to put it outside. Why -- nobody knows for sure, but most likely a mistake. After all it's C++, and by the looks of the rest of the code the author seems like they felt challenged by the language, so they felt like it'd be too much work if / when they realized that the encoding of the message length was misplaced to put it in the right place, and so it continues to this day.
For the record, I implemented a bunch of similar binary formats, eg. AMF, Thrift and BSON. The problem in Protobuf isn't some sort of a theoretical impasse. It's really easy to avoid it, if you give it like an hour-long thinking, before you get to actually writing the code.
Why would it break it? It may make it slightly harder to parse, but since the header also determines the end of the message, anyone parsing the outer message would have a clear understanding that the inner header can be safely ignored as long as the stated outer length has not been matched.
1. Many transports that you might use to transmit a Protobuf already have their own length tracking, making a length prefix redundant. E.g. HTTP has Content-Length. Having two lengths feels wrong and forces you to decide what to do if they don't agree.
2. As others note, a length prefix makes it infeasible to serialize incrementally, since computing the serialized size requires most of the work of actually serializing it.
With that said, TBH the decision was probably not carefully considered, it just evolved that way and the protocol was in wide use in Google before anyone could really change their mind.
In practice, this did turn out to be a frequent source of confusion for users of the library, who often expected that the parser would just know where to stop parsing without them telling it explicitly. Especially when people used the functions that parse from an input stream of some sort, it surprised them that the parser would always consume the entire input rather than stopping at the end of the message. People would write two messages into a file and then find when they went to parse it, only one message would come out, with some weird combination of the data from the two inputs.
Based on that experience, Cap'n Proto chose to go the other way, and define a message format that is explicitly self-delimiting, so the parser does in fact know where the message ends. I think this has proven to be the right choice in practice.
(I maintained protobuf for a while and created Cap'n Proto.)
The problem with writing the length out at the beginning of the message is that you need to know the length before you write it out. For large objects that may cause memory issues/be problematic.
In many cases it works just fine. I doubt any protocol puts a "0" as the length for a dynamic length, but I can see a many-months long technical fight about that particular design decision.
That's not quite right. When a Content-Length is present, it is always correct by definition (it is what determines the body size; there's no way for the body to be a different size). What you probably mean is that you shouldn't ever assume that a Content-Length is available, because an HTTP message can alternatively be encoded using Transfer-Encoding: chunked, in which case the length is not known upfront. That is true, but doesn't change my point: Either Content-Length or chunked transfer encoding serve to delimit the HTTP entity-body, which makes any other in-band delimiter redundant.
What IMO is perfectly ok. Chunking things is a perfectly fine way to allow for incremental communication. Just do it up-front, and you will not need the complexity of 2 different sizing formats.
"Streaming" in gRPC (or Cap'n Proto) involves sending multiple messages. I think this is probably the right approach. Otherwise it seems very hard to use a message as its streaming in as you can never know which fields are not present vs. just haven't arrived yet. But representing a stream as multiple messages leaves it up to the application to decide what it wants in each chunk, which makes sense.
I also fully agree with the "streaming can be done as multiple messages" approach; from the discussion here, it sounds like there may be some nice use cases where having a length prefix would be prohibitive (e.g. compression being generated on-the-fly), but these don't sound like typical use cases for encoding formats intended to be used generally; if anything, I'd expect something like a gzip response to be sent back as the entirety of a response (e.g. an HTTP get request for a specific file) rather than a part of a message in some custom protocol using protobuf or something similar.
You cannot do it in Protobuf anyways (you need to allocate memory, remember? and you need to know how much to allocate, so you need to calculate the length anyways, you just throw it away after you calculate it, fun, fun fun!).
The way it works in real life if, say, you want to serialize a list of varint is that you'd need some small memory chunk (let's call it staging) where you write individual integers (although, this is a bad idea for long lists, as you'd really want to write multiple elements at once if you have enough elements to justify spawning more threads). So, in this staging area you write those integers, more or less a byte at a time. You know they aren't going to take more than 8 bytes in the worst case, so you can have your staging area be 8 bytes.
Then you need to keep track of how many bytes in total you wrote. And, at this point you may start writing the field with the serialized list (in bytes). The field will contain the length of the list. So, you've already calculated the length even before you started writing. Also, you need to store those varints somewhere before you start writing the field with the list...
Protobuf isn't designed to do streaming. Well, really, it isn't designed at all. Like I wrote elsewhere, it was implemented first and then there was an attempt to describe what was implemented and call that "design". Having implemented several formats (eg. FLV and MP4) that were designed for streaming, I'm very confident Protobuf authors never concerned themselves with this aspect.
Length-prefixing is not a problem for streaming. Hierarchical data is, but even then, you have stuff like SAX (for XML).
The problem with Protobuf and why you cannot stream it is that it allows repetition of the same field, and only the last one counts. So, length-prefixing or not, you cannot stream it, unless you are sure that you don't send any hierarchical data (eg. you are sending a list of integers / floats).
Ah, also, another problem: "default fields". Until you parsed the entire message you don't know if default fields are present and whether they need to be initialized on the receiving end.
This can be avoided by magic number. If length is 0, then message length isn't known.
(Also, using zero as a sentinel is not necessarily a good idea, since it makes zero length messages more difficult. I'd go with -1 or ~0 instead.)
I thought -0 is only something in floating point numbers, not integers, and using floats for the length of a message sounds like a nightmare to me.
This is useful when you don't know the size in advance, or if you compress on demand and want the receiver to start reading while the sender is still compressing.
One example could be a web service where you request dynamic content (like a huge CSV file). The client can start downloading earlier, and the server doesn't need to create a temporary file. The web service will stream the results directly and encoding it in chunks.
More accurately speaking gzip (and many other compressed file formats) has the file size information, but that information should (or can, for others) be appended after the data. Protobuf doesn't have any such information, so a better analogue would be the DEFLATE compressed bytestream format.
[1] If you ever have to design one, make sure that reading the first byte is enough to determine the number of subsequent length bytes.
Would it ever be an actual bottleneck though? If it's not actually impeding throughput, I feel like this is more of an aesthetic argument than a technical one, and one where I'd happily sacrifice aesthetics to make the code simpler.
> some message can exceed 2^32 bytes
Fair enough, but that just makes the question "would 8 bytes per message ever actually be a bottleneck", which I'm still not convinced would ever be the case
Anything more requires multiple IPCs, with lots of expensive context switches.
Wasting even one precious byte on a pointless header would absolutely be an issue in this environment.
Not necessarily. Can you really trust the length given from a message? Couldn't a malicious sender put some fake length to fool around with memory allocation?
I was under the impression that something like this caused Heartbleed (to use one example):
When receiving a message, if the user gives you a wrong length, you'll simply fail in parsing their message. Of course, it is up to you to protect against DOS attacks (like someone sending you a 5 TB message, or at least a message that claims it is 5TB) - but that is necessary regardless of whether they tell you the size ahead of time or not.
With heartbleed, a user sent a message saying "please send me a 5MB large hearbeat message", and OpenSSL would send a 5MB re-used buffer, of which only the first few bytes were overwritten with the hearbeat content. The rest of the bytes were whatever remained out of a previous use of the buffer, which could be negotiated keys etc.
My guess is that Protobuf was first implemented then designed. And by the time it was designed, the designer felt too lazy to do anything about the top-level message's length. There are plenty of other technical bloopers that indicate lack of up-front design, so this wouldn't be very unlikely.
EDIT: I'm also still not any more convinced that four bytes per message would ever be a bottleneck for any general purpose protocol, but I'd be curious to hear of a case where that would actually be an issue.
I don't think so. The question of whether you trust the length indication to be correct (you almost certainly shouldn't) seems to me to be independent of whether the length indication comes from inside the message or from some outside wrapper.
My question in the beginning of this thread was intended to be specifically about general purpose formats like protobuf; I think relying on the semantics of TCP or something like that might be a good choose for a bespoke protocol, but it doesn't seem like a great idea for something expected to be used in a wide variety of cases.
One way to make streaming work is just to allow the length value to be bigger than needed and add a padding scheme at the end of the message. This is overhead free in terms of processing time since fields must be decoded sequentially anyway.
In my experience, protobufs are often streamed, especially in the cases where performance matters.
A varint length field prepended to protobuf messages (sent over a reliable transport, such as TCP) seems sane.
Most protocols and serialization formats already define a form of length-prefixed framing; requiring that a protobuf payload also carry such a header would simply be a waste of bytes.
Additionally, it ensures that protobuf can be serialized and streamed without first computing the payload length, which would require serializing the entire message first.
The field prefix byte in Protobuf doesn't really encode "tag and type" as stated in the article, it encodes tag and size (whether the field is fixed 64bit, fixed 32bit, varint, or variable size)
This is pretty self evident when you look at how submessages are encoded the same way as strings, both are just arbitrary variable length blobs.
You cannot reliable determine from a Protobuf message whether a field is an integer, a double, a bool, or an enum without the schema.
Protobufs is a TLV format that just happens to have a compact binary encoding.
I thought this was common and well-known, but apparently not.
> WHY? Oh, why?!
> And yes, it turns out that other people have noticed this anomaly. It's screwed up encoding and decoding in their projects, unsurprisingly. We found a (still-open) bug report from 2018, among others
If anyone knows which library/language these issues the author is talking about are in, please tell us. I'd like to avoid that library if possible
https://github.com/protobufjs/protobuf.js/issues/987 maybe (based on "And yes, that string in this post is entirely deliberate").
Prepending the message with the length means the message is length-limited. Seems standard practice here.
You don’t conflate framing with payload by emitting invalid non-standard framed data from a “protobuf” encoder. They’re separate concerns and need to remain that way.
> you skip the "helper" function that's breaking things.
Yea ok, I'm just going to assume this helper function added framing unless told otherwise. Where in this post did you even read that framing and payload data were conflated (not to mention that there are better protocols that include framing metadata).
The result is not protobuf and doesn’t claim to be.
Looking at my employer’s protobuf runtime for our major programming language, we don’t even support it.
If your implementation doesn't support something in the reference implementation, that seems like your problem, not anyone else's.
[1] https://protobuf.dev/reference/csharp/api-docs/class/google/...
Because it’s not part of the protobuf specification, not part of a valid protobuf message, and framing is a transport/file format concern and should not be performed by default.
If someone is expecting to receive a protobuf-encoded payload, it must not include a framing header.
> If your implementation doesn't support something in the reference implementation, that seems like your problem, not anyone else's.
Someone else’s failure to follow the spec is not our problem.
https://protobuf.dev/reference/csharp/api-docs/class/google/...
...and thus it was added.
We finally had to get down to individual bytes from the network dump to try to sort it out.
The perils of abstraction strike again. Chances are that if you had just written the code to directly send and receive the data you wanted, since you control both ends and know exactly what the bytes will be, you'd never have run into this.
Instead you would have run into different problems..
On the other side, if they wrote multiple projects and never noticed the behavior, how bad can it be ?
It would be interesting if that extra header optimized the processing a lot, pushing other libraries to have it as an option.
Many projects will choose a standard encoding to give them language independence, but start by using the same language and libraries on both ends of the pipe. Therefore,you might not notice the library is incompatible until quite late in a project's development when you try to replace a component with a seemingly compatible, alternative implementation.
https://github.com/protocolbuffers/protobuf/tree/main/confor...
Seems like a good idea for protocols in general to have an official test suite, as a way to address this problem
Framing appears higher up the stack, as an RPC transport, or structured storage like the recordio format referenced by the author. The article sounds like the client expected application/protobuf but the server sent application/custom-protobuf-framing.
Originally OSC was intended to be "transport independent" and in practice only used for UDP transmission, so packetization was left as a problem for implementations to figure out. Cue the predictable problem of different incompatible ideas.
In the specification update it was suggested to use SLIP encoding for packetization, but prior to this the library had already implemented exactly this kind of length-prefix encoding. So now the library allows to select one or the other for sending, and for backwards compatibility on reception the library tries to sniff which packetization protocol is in use. (Fortunately OSC has a predictable first character, and 99% of the time this does not line up with the first byte of the length prefix, so when it starts with that character, the software switches to SLIP mode and scans for the end-of-packet code, otherwise it assumes there is a length prefix.)
It's not ideal, but it works. But it would have been simpler if the original specification had mentioned stream-based transmission. I still get the occasional question about "what are these extra bytes" at the beginning of the message when TCP is used.
Overall I find that a lot of standards have had to be hacked a bit to support packetized transmission, including JSON [0], etc. It's an odd thing to leave out, but admittedly it is indeed a transport problem and most often not considered part of the "file format". So I don't know what the best solution is, but I guess it would be good if data format specifications say something about how to handle multiple serialized instances in a stream if no other standards apply.
Interestingly YAML seemed to have thought about this with the "---" and "..." symbols, even in version 1.0 of the spec [1].
Length-prefixing an otherwise standard format, by default, kept me confused for some hours.
Drives me nuts!
I can’t wait to get back from vacation to ask if it was this.
Both ways are in the standard, but both ends have to agree. One could argue that is a flaw in protobuf itself, but the problem here was not a non-standard implementation.
The problem here was lack of experience and hasty finger pointing.
So, here are just few things that made me decide once and forever never to touch this format:
* It cannot do streaming. It pretends that it can, but it cannot. The problem is that things that should be interpreted as hash-tables or call them "structs" don't require that key-value pairs have unique keys. Also, the last duplicate wins. So, unless you finished parsing all the pairs you cannot call any handlers / construct the "struct" because you don't know if you have the right values for it.
* The grammar is written by someone who... maaaay have seen a grammar... once... long time ago. It makes absolutely no sense. It was written after this mess was somehow implemented, but was never really checked. It's pure nonsense and nobody wants to fix it because nobody really knows how it's supposed to work, nor would they know how to encode in any grammar the actual behavior of the parser.
* Some details like default values for fields that aren't sent or bad (ambiguous) syntax that mixes package names and namespaces...
* Unnecessary constraints on field names that are motivated by unnecessary functionality to translate Protobuf to JSON.
* (C++-specific) the idea of adding messages together is the pinnacle of first-year C++ programmers: they must override a very commonly used operator in a way that breaks every contract that operator makes and for no gains except to confuse and to inconvenience everyone using this operator (usually, accidentally).
---
So, not surprisingly, there's another "feature" in Protobuf that needn't exist, shouldn't exist, and probably most implementations don't implement, but yee-haw! Someone did add it.
My understanding here is that the underlying problem is as follows: "messages" (the composite unit in Protobuf) don't encode their length at top-level. That's idiotic, but that's how it is. So, anyone who wants to send a sequence of messages needs to invent a tiny little bit on top of Protobuf format to tell the other end how to separate the incoming messages. Different Protobuf implementations do this differently. Some use the same encoding as Protobuf uses for varint, others use fixed-length int, some send a whole special header which, beside other things, contains the length information...
My guess is that what happened is that OP and his/her friend chanced on two implementations that didn't agree on how to separate messages sent in sequence. Quite possibly, one implementation wasn't even designed to send messages in sequence, and so had no mechanism for separating them (possibly implying that however uses the library should implement that) while the other one had some mechanism in place. -- I've been there with eg. Scala and C sending messages to each other. This is probably more common that just Scala + C.