The Internet is running in debug mode (2014)
java-is-the-new-c.blogspot.com
java-is-the-new-c.blogspot.com
Pro-text: Human-readability saves immeasurable human developer time!
Pro-binary: Text parsing wastes a lot of very measurable machine time!
Problem is, half of this argument is based entirely on anecdotal evidence and gut feelings. We really have no idea how much developer time is saved by having messages be human-readable. You will find smart people who believe very strongly both ways. In a seemingly very high fraction of these cases, it seems that people are really just taking the practice they are more comfortable with (because they've used it more) and rationalizing an argument for it because it makes them feel good about their choices. When hard evidence is lacking, confirmation bias, unfortunately, takes over.
As the author of Cap'n Proto and former maintainer of Protobufs, I obviously come down pro-binary... but I won't bore you with my argument as I don't really have any hard facts either.
That's been a big win for HTTP. It's easy to imagine the original HTTP authors thinking that 24 bits is plenty of space for file size (anything bigger should use FTP) or using seconds since 1970 in 32 bits instead of text dates.
Those are a bit obvious, but a developer could easily put in something like a 255 chunk limit without thinking too hard about it.
Even John Carmack accidentally screwed over OpenGL driver developers by copying the GL_EXTENSIONS list into a 1000 char buffer in GL Quake (that's way too small).
A binary format makes hard limits too easy.
Sure they do. Where do you think Y2K came from? RFC 822 said "years are two digits" and people wrote parsers to that spec. Even now the official rule (AFAIK) is two or four digits, not an arbitrary number of digits.
And even when the spec has no limit, implementations often intrude. For a more current example, JSON has no "natural" numeric size limits, but interoperability of integers greater than 2^53 is problematic because some implementations assume double-precision floating point while others can parse larger numbers into a 64-bit integer.
Limits are a part of protocol design; a good protocol should make them clearly defined, suitably sized, and feasible to change.
For example, try transmitting 3↑↑↑3 as an integer over HTTP.
Choosing a too-small fixed integer size is one way that a format can screw up and need correction -- one which happens to be (somewhat) exclusive to binary formats. But there are plenty of ways text formats can screw up too, like trying to pack fundamentally structured data into an ad-hoc hard-to-parse string format in order to make it look more pleasant to humans (I'm looking at you, MIME type specifications), or failing to use consistent parsing rules (cookie and user-agent headers totally flout otherwise-consistent HTTP grammar rules, confounding well-written parsers).
If a fairly recent tool is the only way to get binary formats right, then it's entirely understandable that they've fallen out of use.
Protobufs was open sourced in 2008 (after being in use inside Google since around 2001). I suspect other binary formats prior to that had a solution for compatibility as well, but I haven't done a survey.
> You will find smart people who believe very strongly
> both ways. In a seemingly very high fraction of these
> cases, it seems that people are really just taking the
> practice they are more comfortable with
I think you've just solved the Internet, and every programming holy war ever. The irony of course being that tech folk hold themselves up as being particularly rational and immune to this kind of thing.It's a strong indicator that a problem is considerably less important than people intuitively feel it is. Likewise, a great way to tell that a news source is genuinely unbiased is both sides of an argument complaining of its bias.
One of the few (only?) things to love about Go is that a large portion of dull holy war decisions have been premade for you. Generally (imho) wrongly, but who cares? They've been made, and we can get on with solving problems effectively.
HTTP is unusual in that it is highly structured now in spite of having started life as highly unstructured. In addition, once people start using https more and more, we're going to have to use binary decoders on HTTP streams anyhow. So, binary http is really fait accompli, we might as well standardize it.
However, I have a very strong aversion to binary formats for anything persistent that has a lifetime in the range of year+. I have had to grovel through FAR too many binary formats with a hex editor because the program that created them no longer runs because of OS, language, or computer upgrade. This is especially bad when closed-source programs are involved (dredging through SPICE circuit simulator data formats produced by FORTRAN programs of yesteryear is an exercise in pain).
It's true that binary formats present more challenges to "digital archaeology" than text-based formats, but I don't think this is the most important thing to optimize for. Moreover, tools for reading Protobufs are not likely to go away any time soon -- not any more likely than tools for reading MySQL databases or your preferred compression format or your filesystem, all of which are also binary. (Cap'n Proto is perhaps not as deeply-rooted yet, but I'm working on that...)
Of course, again, this is a judgment call and there's no way I can prove that my judgment is "correct".
Internally (between services) the format should be binary. It's important to have appropriate tooling support in order to be able to debug production service communications if necessary.
For development purposes it would be nice to be able to switch to a text format (i.e. JSON) for debugging.
Since JSON is a very convenient format for the web, it would be great to have a library that is compatible to both modes.
I'm currently evaluating Flatbuffers vs. Cap'n Proto. I think that for my purpose they don't differ much. But Flatbuffers implements JSON reading/writing. I like that Cap'n Proto has forwards- and backwards-compatibility. So I would be interested in your plans on implementing the JSON codec support (mentioned on the Cap'n Proto Roadmap).
When you as a developer look at JSON messages, and you see endless walls of irrelevant text scroll by, you're disgusted. This disgust drives minimalism. With a machine-centric format, the excuse quickly becomes "oh, this is intended for machines anyway, so who cares". If you can access something through tools only, the bloat becomes hidden and is encouraged to grow.
At that point, you'll start seeing articles asserting that, sure, the new super-efficient binblob messages are 10x-20x as large as JSON used to be, but look at all the things we gained, like automatic protocol negotiation, contracts, actual serialized objects. Any of these sounds reasonable at first but in reality will only benefit tool vendors in a vicious feedback cycle where the format slowly evolves itself to death.
I take that 3-5x overhead of parsing JSON any time over the non-human-readable alternatives. That doesn't mean it's the right choice for all protocols. But it's a reasonable default for a lot of systems.
I never said it was. I apologize if it wasn't clear from my comment, but it was centered solely on the aspect of human readability. You're not wrong about it being speculation though, I just don't see the harm since we're talking about a speculative source article in the first place.
> Can you provide some examples of actual binary protocols that underwent this transformation?
My entire comment was addressing an example where I felt a format had degenerated because it transitioned from a human-readable to a decidedly machine-readable form. My criticism is not based on the notion that binaries are bad in of themselves.
XML was never intended as a memory-efficient format...
I wonder if it was assumed during the creation of XML that it would usually be compressed. The largely symmetric open and close tags compress quite well...* Is it perfect? No. Is any protocol? No. Would a binary format be better? Very likely, no.
We've already agreed on a binary protocol: UTF-8 (previously it was ASCII.) But we've also built redundancy into it for the humans to make sense of it with their high-level brains. Instead of a single byte representing an HTTP header, we use a string of bytes. Now the human involved can tap the wire and watch the request in real-time without processing anything.
Now, if you'd like to remove redundancy without the need for a compression library, we'll just need to agree on shortening those strings. And we'll need a new diagnostic/parsing tool for each [binary] protocol that's invented -- unless you can convince the grep/sed/awk developers to add every protocol to their tools. Or maybe we could all agree on a single binary encoding for every potential combination of strings; something like an index into a dictionary. It might be better (i.e. higher compression ratios) if we let the computer decide on the dictionary for each message.
Do you see where this is headed?
1 - This, of course, is only the case until the machines can accurately gauge human intent and respond appropriately, preventing us from making mistakes to begin with.
> Standardize on some simple encodings (pure binary, self describing binary ("binary json"), textual)
Maybe like gzip [1], hpack [2], bson, or others?
I realize the point he's making about doing unnecessary work, but there's also a reason we haven't expanded human language past written characters or spoken syllables. It's efficient for our brains, and for preserving and transmitting knowledge.
There's just no way to create a binary format (character encodings aside) that can encompass all the possible ideas that can be communicated. Instead, the common text protocols eventually get optimized into binary (HTTP/2) without compromising the ability to express the rest.
[1] https://www.ietf.org/rfc/rfc1952.txt
[2] https://tools.ietf.org/html/draft-ietf-httpbis-header-compre...
It's funny that you mention IP and TCP. I's sad that we've degraded the end-to-end principle those protocols embodied. Today, it's not really practical to use IP protocol numbers other than 6 (TCP), 17 (UDP), and 1 (ICMP) and their IPv6 equivalents. Middleboxes, out of misplaced caution, reject packets that look unusual. We can't even use ECN. To work around this problem, we run everything over TCP and UDP.
That doesn't even work, though, because other middleboxes, working one level up the stack, reject TCP on anything but ports 80 and 433.
So now we run everything over multiplexed encrypted connections on one TCP port on one IP protocol. That's just silly.
The low-level protocol situation is like a fossilized idiom in human language. In English, we can that something "wreaks havoc", but what about "wreaks happiness"? Nobody understands "wreaks" anymore. It's no longer a productive rule. Likewise, IP and (arguably) TCP protocol details aren't productive either, except in specialized cases, or on closed networks.
The internet isn't just running in debug mode. It's also flinging around lots of bytes that do nothing other than reflect past aspects of its evolution, like an embryo's tail and gills.
Anyone can look at a web page and learn how it works. Anyone can copy bits and pieces of code from various places and put it together into a web page of their own. Most "web programmers" manage to make a living without ever having to learn a complicated protocol or trying to figure out what a long string of hexadecimal digits means. Easy to learn = more casual tinkerers = more people learning how to code, at least at a basic level.
You could design perfectly efficient protocols and encoding formats, but if people don't use it, what's the point?
HTTP/2 is a nice compromise. The web server and browser abstract away all of the binary layers, so most programmers only need to care about human-readable text.
Not so true anymore. It's true that HTML/CSS are easy to inspect on a browser, but the meat of most websites today (JavaScript) is heavily minified to the point that you'd be hard-pressed to call it plain text.
So even if not all websites are in debug mode, there are incentives to make it easy to get them into debug mode.
Its an interesting claim, and it is certainly true that encoding numeric information into UTF8 consumes CPU cycles. But what isn't quite so clear is "What percentage of packet latency is dedicated to encoding and decoding packets?"
Back in the old days I was the ONC RPC architect for Sun and we spent a lot of time on "RPCL" (RPC Language) which was a way to describe a protocol textually and then compile that description into library calls into XDR (the eXtensible Data Representation). We did that because you burned a lot of CPU time trying to parse a C structure out of the network, and more importantly the way in which it was represented in memory was an artifact of the computer architecture (bit endianness, did structures get packed or were they word addressable, etc) XDR solved all of those problems by putting data on the network in a canonical format, and local libraries could always convert from the canonical format into the local format correctly.
That actually works quite well. It almost became the standard way to do things on the Internet, but politics got in the way. The big argument was that if you converted things into big-endian form on the network, then a little-endian processor had to convert to send and convert to receive, but a big-endian processor got a free pass without "painful" conversion steps.
Later, rather than converting big endian to little endian people just convert to text (which has the same effect of a canonical form) but it hides the religious argument behind the "hey its just text, we all know how to parse text right?" sort of abstraction. At least then it penalized everyone equally.
But the truth, which came out in the RPC wars, and is even more true today, is that you have to burn a few billion CPU instructions to have any impact at all on latency. That is because computers are so much faster and the network? While faster isn't a million times faster, it isn't really is just barely, and only on a good day a hundred times faster than it was back in 1985. What that means in principle is that if it takes 1 uS or 30 uS to queue your packet for the internet it doesn't even show up on the carve out of the 5,000 uS it takes to send a small packet from here to there, or the 200,000 uS it more typically takes.
If you're a supercomputer sending data around to simulate fluid dynamics, that stuff adds up. If you're sending ajax calls from here to there, not so much.
I just recently used ONC+ for controlling a low spec, low-as-possible-latency, embedded Linux controller. The ridiculously low memory footprint and CPU usage, compared to all of the others, along with the whole 5 minutes it took me to write the server and client stubs with rpcgen made it an easy winner. So I thank you. :)
We also tend to forget, or never learn, what was discovered before our careers began.
Not only hardware is cheaper, but is much more easily replaceable, have a well defined behavior and usually don't go nuts and need to take prescription drugs. Also they mostly don't have a life and family outside work.
This is why technology usually converges to serve the humans, not the machines.
Or else we would all be writing assembly code instead of garbage collected, duck typed and compiler aided languages.
Whole Web is ridiculously inefficient. Even in 90's, there were better architectures [3] to choose from. It's unsurprising that "Web" companies such as Facebook have moved back to Internet apps where possible (esp mobile) and often avoid vanilla Internet/Web protocols within their datacenters. There's better stuff in every area. Here's to hoping that more of it gets mainstreamed.
[1] https://en.wikipedia.org/wiki/External_Data_Representation
[2] ftp://ftp.cis.upenn.edu/pub/cis700/public_html/papers/Franz97b.pdf
Sure, HTML, CSS and JS are "text", but it's usually optimized, compressed text, and pretty soon we'll be shipping that around in HTTP2.
People optimize what matters. Look at the HTML source for this page. It's mostly user content. More efficient packaging for the HTML structure wouldn't make a measurable difference in load time.
https://developers.google.com/protocol-buffers/
Unfortunately it was made public only after JSON had already become "the XML replacement".
I used to be obsessed with performance and refused to write in anything other than assembly. When I started working with teams of other people I realize there is efficiency for the computer and there is efficiency for the developer and the company. The latter 2 trump efficiency of the computer pretty much every time.
I'm more than happy to sacrifice some efficiency if it means getting the job done faster, cranking out more features, greater compatibility, more leverage with existing tools / standards, and so on.
If you want to talk about efficiency though. Why don't we get rid of time zones and daylight savings time?
Yep. <img src="whatever.jpg"> is indeed all text. and zip would have turned those 25 bytes into 191. And then a 300KB highly-compressed binary data file would be transmitted.
aka "Penny-wise and pound-foolish."
I've considered going back to that in custom designs. It's just that Galois built a high assurance ASN.1 parser, INRIA has a verified parser generator, and ZeroMQ is pretty solid. Too many neat things to consider these days... haha
JSON is easier to use, debug, and is more efficient than XML.
WebAssembly is going to reduce data transfer sizes and load times while increasing developer productivity as the ecosystem and tools surrounding it expand.
That's being a little too optimistic. And assuming that does happen, it's going to take years before it surpasses the current toolset.
Suppose HTML had been a binary format. Would it have gotten anywhere near as far?
CBOR is great, it's similar to JSON & already an RFC http://cbor.io
Written by someone with "Java" in the URL. Why did I even click.
Perhaps you feel there's no point to doing that either, but then the only HN-appropriate thing to do would be to just not comment. Posting empty swipes doesn't help anyone, and breaks the site guidelines.