For those of you interested, I suggest using Google's protobufs or even full blown gRPC for such endeavors.
See discussion about gRPC here: https://news.ycombinator.com/item?id=12344995
For those of you interested, I suggest using Google's protobufs or even full blown gRPC for such endeavors.
See discussion about gRPC here: https://news.ycombinator.com/item?id=12344995
Before looking for something 'better' than JSON, build a basic prototype of your app and see if your application actually needs to use something else. The data in most applications is so small and will parse so easily that actually measuring the CPU utilisation will be quite hard. 99.99% of applications aren't moving data around at the sort of speed where the format really matters.
Protocol buffers are very clever and really cool but using them where you don't need them is a huge waste of effort.
As for the lack of a schema that's not true. There is a schema. What's in the file is defined somewhere. For applications that consume JSON that's in your code rather than the file format.
Plus, the byte count over the wire is much smaller - something that matters even in today's world of gigabit ethernet.
Are you speaking from personal experience? FWIW it doesn't match mine.
We benchmarked one of our apps, written in Go, which uses JSON extensively both for the API and for internal storage, and we found that around 90% of the CPU time is spent serializing and deserializing JSON.
The JSON code itself is a tiny portion of any app, so it's one of those rare optimization cases where a small change can yield significant speed-ups for the entire app.
On the other side, you can produce JSON string directly, saving the need to produce some intermediate representation in memory.
You should be able to saturate 10gbit line on a multicore server easilly with this approach.
That implies that JSON is somehow the gold standard of serialization formats. I don't see it that way. The fact that JSON is ubiquitous in some areas doesn't mean that it's the gold standard for all other uses.
In my experience, protobufs are easy to work with, concise, safe, and efficient. There's never any ambiguity to them, and both creating and parsing messages is fast and easy. That doesn't mean that there aren't cases in which JSONs aren't a better choice, for instance when working with browsers, or needing a text based format for debugging or inspection purposes.
> Protocol buffers are very clever and really cool but using them where you don't need them is a huge waste of effort.
In my experience protobufs require very little effort, with most of the effort being put into defining the schema. It's pretty frictionless afterwards with one notable exception - when you want to manually inspect a message. That's where JSON wins hands down.
I'm not suggesting JSON is the best format, because that depends on the project you're working on. I'm saying it's good enough for most tasks, and has ubituous support, and it's easy to work with. There are better formats for particular use cases but if your application has requirements that can work with JSON it probably isn't worth finding something else.
Contrast with JSON where every decent language ecosystem has a parser readily available. That's what makes it the "gold standard" for me.
No affiliation, just satisfied user.
JSON-schema implementations are also rather widely available.
In performance-sensitive machine-to-machine applications, I agree, better to use someting compiled, but for general-purpose interoperation and archive data storage (where you cannot know when and with what tool you'll want to open it), I've yet to see a compelling alternative.
ProtoBuf depend on the schema definitions and encoder/decoder generation. Without the schema, you cannot read the format (theoretically), if you want your data to be read easily, JSON or CBOR or MsgPack are your best options.
There are few things as slow as non-vectorized MATLAB code. MATLAB's regex engine is pretty decent, actually, and supports quite a few fancy tricks. For example, I use the regex engine to find the start/end of all numbers and strings ahead of time, and for replacing escape sequences (even unicode escapes!) with their literal characters. The former is particularly important because otherwise, you'd have to manually iterate over characters until you find the end of strings or numbers, which is unacceptably slow in MATLAB.
Honestly, that statement applies to pretty much any part of MATLAB unless the behaviour is nailed down by existing mathematical or engineering specifications - and I'm talking as someone that likes MATLAB overall. Mathworks tries a bit to polish and paint it as a coherent whole, but it's patchwork of things built upon and stuck to the sides of the existing structure, optimizing towards a particular kind of "pragmatism" than towards any kind of elegance.
There's also the ability to easily and explicitly encode data that doesn't fit to JSON's limited types. A non-exhaustive example includes: sets, hash tables that don't use string keys, tuples, imaginary numbers, fractional numbers, and precision decimal numbers.
There's also the fact that XML explicitly supports streaming interpolation of documents, something which can speed up interactivity quite dramatically.
Readable plain-text protocols remove a constant overhead from all debugging/introspection tasks. If you use, say protobufs, you'll still have to convert back into strings to see what was in individual messages. This makes reading, for example, logfiles or wireshark captures with "grep" more difficult. Sure, there are tools that you can throw before "| grep" that convert to text, but that's one more piece of cognitive/debugging overhead you have to remember, and one more thing that new developers have to know to do if they're used to the standard unix/text-oriented way.
That benefit applies even if your messages (or logs of messages) are stored/viewed in a database that speaks your binary protocol and provides tools to search/see the data as text. Whether or not you have something like that, it will always (always) be routinely necessary, or at least incredibly helpful to productivity, to "print out or search the textual content of messages sent/received by my code" when developing and testing apps that communicate over they network. A text protocol has a huge advantage in that department--or rather, it has a very small advantage that is present in a very large number of development/testing tasks.
Also, readable plain text is key. XML dialects, or e.g. Thrift's JSON format are plain text, but are not terribly readable. This is subjective, to be sure, but makes a big difference in practice.
There are great reasons to use binary protocols, but any comparison that leaves out the benefits of text content that is readable "for free" without using per-protocol tools is lacking.
As others have pointed out, text-based serialization has significant overhead. To save a few hours of debugging, you're willing to spend gigabytes of bandwidth, an uncountable number of CPU cycles and memory, and all the associated electricity and hosting costs over the lifetime of your service? Penny wise and pound foolish, to say the least.
Text-based serialization formats are just plain stupid for any type of service that handles over 100 requests per second. A good developer is not hampered in the slightest by the fact that their serialization format isn't text.
Would I switch to a more computationally expensive format to save a few hours debugging each time someone makes substantial changes to my application's message format? Very possibly.
Would I switch to a more computationally expensive format to prevent bugs (the easier it is to view, or remember how to view, what your code is sending/receiving, the less likely it is that lazy people will skip bug-hunting/QA steps that would show it)? Almost certainly.
Should you use a binary format if you're sending tons of data per second? Well, it depends. It's like microservices: if you have the tooling to make dealing with that format from the dev/debug/tracing side a breeze, go for it; using binary provides substantial savings, as you point out. If you don't have the tooling/time to make it as easy as text to introspect, and I mean that in the most absolute way, there might still be persuasive reasons to switch, but know that if you do, you are also buying a lot of "debuggability debt".
(I know there are good reasons for this involving the effort required and maintenance etc. it just bugs me.)
In Windows, PowerShell is a very good example of a low-ish-level (composable commandline apps and snippets) programming environment that isn't shackled by that paradigm.
Please promote this as hard as you can. Then, us old guys getting aged-out of industry will be able to get jobs again.
Make “ASN.2” or something, update the fundamental data types to modern standards (think of what changed from NeXTSTEP to macOS X; they changed from TIFF to PNG, from PostScript to PDF, etc. but kept the overall architecture the same). Consolidate and simplify the many options (but do not eliminate options entirely). And call it… JASN or something, which sounds like both JSON and ASN, and make up some acronym where the “J” stand for “Joint” (like in JPEG, which stands for Joint Photographic Experts Group).
And, maybe most importantly, make freely licensed tooling available for both converting to and from ASN.1 and JSON, but also jgrep, jless, jcat, etc. (similar to zgrep, zless and zcat, etc.) Make these packages available for all common operating systems, readily installable using that plaform’s usual mechanism.
Ditto. And if you still prefer the schema-free approach of JSON but with faster encoding and decoding support across all common programming languages, take a look at MessagePack.
Unlike JSON, MessagePack is optimized for speed and encodes common data types efficiently. Unlike Protobuf/gRPC, it requires no centralized schema.
MessagePack: https://msgpack.org/
I say that a bit tongue in cheek, but really: BSON has a monstrous amount of basic data types, some of which make very little sense and are very difficult to round-trip through other formats; and MessagePack had crippling basic sanity issues with what-are-bytes-vs-strings in its early history and that's something you simply can't recover from without declaring game over and making a new spec.
CBOR has neither of these issues. It's almost completely isomorphic to json, plus supports byte arrays. It's very nice.
I haven't played with flatbuffers, but it probably supports this as well, given its design. Are there other binary protocols that have first-class (i.e. present with a good API in the client libraries for all major languages) support for lazy/partial deserialization?
EDIT: I am happy to see that it handles date-time. Not really my favorite part of the non-spec of JSON on that front...
They use CBOR as the underlying serialization, and have a deep interest in canonicalization and applications of hashing.
The json schema is standardized and available in most languages and platforms.
It's like saying ruby isn't the only ruby language because it isn't a formal standard. So what?
I'd rather use Protobufs or Flatbuffers for specification and then rely on their built-in JSON serialisations.
Remember that the more modern binary protocols like protobuf/etc. don't serve the same niche as JSON schema at all: JSON schema validates external documents; i.e. you take a schema and make sure a document complies with it. Protobuf/flatbuffers/etc. validate internal documents, as they are encoded or decoded. If that's all you need, then the schema definitions for those products will be much friendlier than JSON-schema, but that's comparing non-equivalent things.
Another, minor quibble is that a lot of the schema-validation code in e.g. protobuf or Thrift is often generated by the framework, and the generated code can be quite hairy or confusing. There exist code generators for JSON schema as well, but writing a JSON schema by hand isn't comparable to, say, generating code from the Thrift IDL.
Example: when you add a new element to the API response, the recipient needs to update their schema.
Is that correct? If yes, how do you deal with that in a microservice architecture where you have little control over consumers of your API?
You are free to make forward compatible API changes without a deployment dance.
E.g., you can add a field in the service and the caller won't see it until it is prepared to handle it properly.
... it should never be fragile where fragility might cost more than it gains you. The open/closed principle is the right thing to follow in most cases, but not all.
For example, choking on unrecognized additional fields by default in development is often a big time/bug-saver, since it gives you an early "hint" that things might not be speaking the same dialect to each other.
Similarly, if you have a message convention that is heavy on "modifiers" which control irreversible things, sometimes it's better to be fragile rather than count on orchestration systems to work perfectly 100% of the time. Suppose you have a system for buying bicycles with a required field of "bike model", and modifiers of "color", "wheel size", etc. Now let's say your servers all got patched to expect a new modifier of "frame size" . . . except for one server which didn't restart due to a deployment bug, and is still serving traffic. If a client sends that unexpected additional modifier to the "old" server, and it gets ignored, this could result in sending a production order for an expensive unit with the wrong specifications. If not caught, that cost could be compounded by sending the wrong unit to the customer and pissing them off. Now, this is clearly not a problem with the message format, but rather with the orchestration/release system. But bugs in orchestration layers are incredibly common, and baking a bit of fragility into the messaging layer at the right places can help mitigate those bugs' impact.
That's not the case with protobufs. The sender can add new fields, and they will simply be invisible to receivers using the previous version of the message.
You'd need to update the receiver only if you want it to access the new fields, which is what you'd need to regardless of serialization format.
If the issue is debugging, writing a dump tool is a one time effort.