The 7-bit Internet
blog.tabini.ca
blog.tabini.ca
It was also a fun way to learn programming. mIRC at some point added a sockets interface to mIRC script, and after that it was pretty easy to write toy clients for various text-based protocols, doing simple things like checking whether a URL was a 404. I can't say I did anything particularly useful with that, but it was a nice way to learn something about programming.
I think TAOUP sums up a few important points about text protocols [1] and why going binary is not always wise. Obviously the main drive here is money and I'm sure we all have stories where short term money translates later to large unforseen and/or hidden costs.
Bandwidth capacity is growing exponentially so do we really need this, even for the savings ?
I'd rather see people invent clever new ways to do wonders with text streams than go back to opaque and obscure data structures that bring back to me rememberances of old days proprietary protocols.
After all, if our fathers made the choice of text streams at a time when each byte cost much more than now, there may be good reasons.
[1] http://www.catb.org/esr/writings/taoup/html/ch05s01.html
The ease of discoverability of a protocol should not be underrated.
I guess if you've never had that moment where it all clicked for you, just by looking at traffic, the significance of human-readable protocols is easy to underestimate.
But hey, Google says that with binary-compressed headers of sorts, they can fit all their tracking cookies in one TCP-packet and that means you wont feel inconvenienced by the tracking payload they are throwing at you everywhere you go on the internet.
You know what? I want that huge bag of tracking cookies to hurt. If you violate my privacy, that should come at a cost.
Google: You can take your SPDY and keep nice, old HTTP alone.
Length-prefixed fields and direct encoding of binary data go a long way to simplifying implementing a protocol.
The ease of implementation of a simple binary protocol should not be underrated.
Complex + non-discoverable = bad engineering
I wonder if that was the original purpose of the protocol;)
The fact is, on the application layer, text based protocols wiped the floor with binary protocols. The answer is in the little details, which everyone knows is where the devil spends its time. They are not easier to implement, because they're more difficult to debug (and because text parsers are, let's be honest, a solved problem); they are not as readable, because debug tooling is never perfectly implemented and never ever present in every system; they are more efficient, but at this layer the efficiency does not pay off (i.e. the size of the HTTP header leading that 5MB image does not really matter).
One of the comments here pulled a comparison I did not remember any more: SS7 vs SIP. Go and have a look at both protocols. They are both mature, so you don't really need to wait for the fantastic debug tooling to appear. Study both ecosystems and then form your opinion on binary vs text application protocols.
Do you remember the OSI binary protocols as designed back in the day?
ASN.1 as far as the eye could see. OIDs. MIBs. Loosely defined extensibility. Terrible, complex protocol designs.
This had nothing to do with the protocols being binary or text; the protocols themselves were painfully complex, making it difficult and frustrating to write a working implementation, much less a complete and interoperating implementations.
> ... at this layer the efficiency does not pay off (i.e. the size of the HTTP header leading that 5MB image does not really matter).
"I don't care about RTT time, the lack of bidirectional communication, and the inability to inline binary data in the protocol stream" ... said no mobile/wireless/desktop developer, ever.
> One of the comments here pulled a comparison I did not remember any more: SS7 vs SIP.
I've implemented SIP. It was a massive hassle because it was text-based and annoying and difficult to parse. The available SIP libraries tend to be buggy and incomplete, and people actually pay for workable SIP stacks. I'd rather not look at SIP ever again.
That isn't to say SS7 is better, it doesn't appear to be. As far as I can tell, that's because it was badly designed, not because of a flaw inherent in using binary encodings.
+1
http://www.igvita.com/slides/2013/breaking-1s-mobile-barrier... if you want to read about this from a mobile perspective.
Now we build a different tool, let's call it http2net. With HTTP being so prevalent, perhaps it isn't unreasonable to think it becomes part of most OS distributions.
So now you are using a tool (http2net) to view/interact with the new protocol and this is good because it is easy to read and use.
In summary, is binary vs text really the big problem here?
You'd think so, but while it's true we can mostly pretend telnet gives you a raw TCP connection, telnet intercepts some character sequences. It just doesn't interfere much with plain text.
My problem with the hypothetical "http2net" tool is that to this day I still have to deal with systems that doesn't have basic tools like tcpdump, curl or wget. But telnet, nc or some other way of getting a semi-raw tcp connection is pretty much always available.
So I have every expectation that it'd take 30+ years before this "http2net" tool would be available everywhere I'd want it.
Regardless of the issue of negotiated options, most telnet clients goes into a command mode if you enter an escape character, defaulting to ctrl+], unless you explicitly change it or tell it not to with a command line switch. This happens for me regardless of port.
tcpdump and strace also confirms what I thought I knew, namely that it also does LF => CR+LF conversion also when connecting to other ports.
EDIT: Here's a simple demonstration of how it is decidedly not 8-bit clean:
echo -e "test\035help" | telnet www.google.com 80
Not only will this get you the help text for telnet rather than send the unmodified byte stream to Google's unsuspecting web server, but with default options not a single byte will generally get sent over this connection on Linux at least, as the client starts out line buffered and gets put into command mode without getting a line feed.Confirm with tcpdump, or this way if you have strace installed:
echo -e "test\035help" | strace telnet www.google.com 80 2>&1 | grep sendto
(the sendto calls you will get are DNS lookups) echo -e "test\035help" | telnet -E www.google.com 80
will not give any help text output.(I was, in fact referring to telnet option negotiation which will not happen if not using the telnet port. Since the escape character for interactive use is so easily remedied, it did not occur to me to consider it a problem.)
Engineering should be about taking complex things and making them simple and easy (the best engineers do this), on the other hand you have engineers that try to add complexity because they understand something and over engineer it to obscure it on some premature optimization or cool factor. Taking something simple and making it more complex is the epitome of bad engineering.
Cool is useful and useful is simple or at least simple parts. Keep entry to technology simple, just like good games are simple to start, deeper in it gets more difficult. The door should be easy to enter even though the labyrinth might be immense.
The text-based aspect isn't bad in and of itself. It's the crazy moronic rules and pointlessly flexible syntax that makes it bad. It's the fact that "text-based" is often taken to mean "should allow humans to be flexible in writing it", instead of "uses ASCII". Oh, and UTF8 if you're lucky. If you dare want a non-ASCII value that needs to go into a header, you're toast.
HTTP: Comments inside headers, line folding, context-sensitive header value parsing, and on and on. There is zero legitimate need for these things, yet RFC writers cannot seem to help themselves and design the most complicated syntax as possible.
Text-based formats frequently trick people into thinking they're right because it looks easy on the surface to get something to work. Shit, most HTTP clients/servers aren't actually fully compliant.
It opens more possibility for security holes due to potentially ambiguous parsing (where one implementation misses an edge case, but another doesn't). Protocols are not programming languages. They don't need flexible syntax.
There's also crap left over from a time when people read/write these protocols by hand. Idiocy like the IETF preferred date formats "Sun, 06 Nov 1994". Really? Including "Sun", the English 3-letter abbreviation, in a protocol? How is that useful?
HTTP is being used more and more for high-performance work. Take a look at a high-perf text-based parser, and you see all sorts of ugly hacks. Like doing bitwise comparisons on word-sized integers to determine the request method. (Both my personal code and nginx ended up with similar solutions, so it's safe to say it's a common approach for perf.)
Having a proper binary format that can be quickly, unambiguously, and safely parsed is a huge boon for interoperability and performance. The slight detriment for analyzing raw bytes (which, when encrypted makes it all moot) is not worth it. For a popular protocol like HTTP, you're gonna have plenty of tools to properly parse and analyze.
IP, UDP and TCP aren't text based, and I have no problems regularly analyzing them, nor do they seem to have adoption issues. But you can bet your ass if UDP specified port numbers as a flexible text field, you'd find all sorts of fun bugs and implementation issues.
Edit: After skimming the HTTP 2.0 spec, implementing this as a text-based protocol sounds like a nightmare with no benefit. It's not like you're going to write multiple streams out by hand or something (and use some terrible multipart-mime approach, or another "fun" ASCII-delimiting-binary thing).
Though I don't really agree with you that you need a binary format for interop and performance. You'd get the same benefits if people just constrained the text protocols. E.g. Your performance "hack" for HTTP request methods is only a hack because method names can be variable length. We could constrain method names to 4 bytes and make the whole thing cleaner while still getting the text protocol benefits of being able to manually inspect request/responses. Similarly, picking a sane date format like iso-8601 and restricting it further to a specific set of options would do wonders while still being readable/writable for humans.
I do agree with you that once a protocol gets as widespread as HTTP most of the benefit is lost as we get better tools to use anyway, though. Especially with debug tools built into most browsers these days.
The issue is not that non-text protocols are impossible to analyse, but that each new non-text protocol requires new tools to do so. For IP, UDP and TCP we've had decades to create good tools. If a common binary serialisation format was agreed and most protocols stuck to that, it would be much less of a big deal to ditch text-only protocols and rely on client side tools to allow reading or manipulating them as text.
Or, as the langsec crowd claims, protocols don't need flexible syntax, lest they actually become programming languages.
I think it's sad that HTTP/2.0 is primarily just trying to multiplex TCP and the rest seems like pure micro-optimizations (the one exception being server push; I haven't read a lot on that yet). It seems like the wrong layer in which to attack the problem, and is basically a huge 'meh'.
You can have simple text protocols and complex binary protocols, but it's more tempting to make text protocols more complex.
(In this comment, simple/complex is in terms of syntax only)
Google is not the only one who benefits from reduced latency, fast page loads, efficient use of SSL sessions, and server pushed resources.
If anything, the people who benefit the most are those who _can't_ afford massive forward deployed CDN networks, large servers, and fat network pipes.
If the internet should have taught us one thing so far, it is that it's the open technologies which are built to be easy to understand and explore for humans which have driven the net forward.
HTTP 2.0 is massive step backwards in that respect, and the only thing we're getting in return is slightly shorted response times.
You know what? We can deal. We're getting better and better bandwith. We're getting more and more computing power. A few milliseconds doesn't hurt anyone.
Let's not create a shit protocol violating everything the internet was built on, just because we're all of a sudden feeling resource-constrained. Now, of all times.
rolls eyes
Yeah, but the speed of light isn't getting any faster, and little guys can't afford to have their servers everywhere. If HTTP 2.0 reduces the round trip count, then it will make a big difference in download times for small pages.
While we can certainly wait and see whether people will balk and make changes, the only true way to stop it is to replace it.
If you can write a better set of protocols that meet the requirements, do it. If you can't then work with or support those that will.
We have IPv6 connectivity from our colo provider in one location, and a tunnel to our office, but two years in we're still only testing it for this reason - there's never enough time, and it's not yet urgent enough.
Different than the spirit of HTTP, those protocols have little to do with publishing content anymore. They are just more kludges, in the history of kludges, to patch browsers into application platforms.
That mainly benefits the big web monopolists, who require the browser to be the ultimate application platform, where they can track to their hearts content, display unsolicited advertising, and basically extort business to advertise on their channels to remain relevant on the web. Not quite the idea of "information repository" that spawned the web in the first place.
Please. All you need for advertising is plain text, all you need for tracking is cookies.
What is the end-user benefit in performance using HTTP 2.0? The only analysis of SPDY I saw indicated low single digit percentage benefit vs HTTPS. Whoop-dee-do. http://www.guypo.com/technical/not-as-spdy-as-you-thought/
I would highlight that HTTP 2 > SPDY as far as standards go, and the "beta" protocol of SPDY is already at 55% market share including IE 11: http://caniuse.com/spdy
For those who are interested in speed optimization for 2013 and beyond, plus why SPDY's faster-than-SSL descendants are crucial to the web's success, have a look at http://www.igvita.com/slides/2013/breaking-1s-mobile-barrier... -- I believe there's a video out there too somewhere. Edit: Video link is on first page of the slides.
What's interesting too is that this talk focuses as much on what's achievable without SPDY as what might be with, down the road. We really need faster SSL negotiation for that first time connection cost.
By making things more efficient, HTTP 2.0 is really going to make startups spend more money. Oh wait, no, the opposite thing.
EDIT: It's not the matter of implementation anyways, it's that something important only for Google is proposed as a standard for everyone, which is insane.
More efficient web servers will, if anything, be even better for startups than they are for Google, since they lower the barrier to entry.
You win some, you lose some. The benefit of cutting down on some infrastructure should be weighed against the mind-boggling complexity of HTTP 2.0 and the added difficulty in debugging.
Like scripting languages, a text-based protocol doesn't just make it easier for you to get your hands dirty: it practically begs you to. The value of that sort of encouragement to the adoption of a global standard should not be underestimated.
>It wasn’t long, however, before I realized the true genius behind this decision.
This reminded me very much of a talk by Jonathan Blow where he's talking about how he "nerd raged" while reading the Doom source code. He found a section that wasn't optimized, but in fact not optimizing it had so many other benefits (besides, the performance gain would have been negligible). Starts at 14 minutes: http://the-witness.net/news/2011/06/how-to-program-independe...
cake:~ mali$ telnet google.com 80
Trying 74.125.235.8...
Connected to google.com.
Escape character is '^]'.
GET / HTTP/1.1
.......
HTTP/1.1 200 OK
Date: Wed, 10 Jul 2013 22:25:02 GMT
Expires: -1
Cache-Control: private, max-age=0
Content-Type: text/html; charset=ISO-8859-1
At which point they go "oh cool" and go do something else. For the rest of us who use curl, wget, requests, Chrome / Firefox developer tools every day (who are we kidding, it's everyone, you liars! :P), the binary transformation would be transparent.Hell, if you're going for pure cool-factor, how is pulling out your hex editor less cool? But in reality, you'd never do this.
For a non-standardized and obscure protocol where tooling would likely be lacking, I can see why human readability is a good idea. But we're talking about the very protocol that makes up the fabric of the internet. Seriously, why?
Give me one good reason.
You've never needed to test an SMTP connection to see what the rejection message was on the remote server (when a user can't get you the bounced message you require)?
You never wanted to see if an SSH port was open and what version was running?
Telnet is available in every router and firewall I have, I can't install curl onto a router to generate a request from a remote network, and I'll never see wget there either.
I do this stuff every day as part of my job, telnet is the go-to, the other utilities are fine, but they usually mask what I'm really looking for anyway, if they are even available on the platform I am using in the first place.
Text protocols, on the other hand, require writing a parser, dealing with encoding back and forth between string representations and binary data, handling line delimiters, etc.
I'll take binary protocols any day of the week. Any cost they incur in not being human readable is offset by the value of them being so easy to implement.
Slowing down everyone forever just to ease telnet debugging is a misdirected optimization.
It's largely irrelevant to me if there are tools out the wazoo to work with some binary protocol if I'm unable to run that tool everywhere.
And there's a huge range between advocating arbitrarily complex and flexible text protocols vs. binary protocols. You can "easily" do text protocols that are picky about field lengths and that use formats that can be parsed much faster than the more complex protocols.
If your protocol is using a small enough, regular enough grammar, it'd also be fairly trivial to allow a binary serialisation of requests or responses as an option without much extra overhead. E.g. start client connections with a word indicating it wants to "switch on" binary and length prefix any variable length fields instead of relying on an end of field marker, for example. (Or make human clients type out a word to switch to the text serialisation).
But very few protocols are so affected by latency in request/response exchanges that binary vs. text is a huge deal. For HTTP moving to a pure binary protocol might make sense because of how heavily we depend on it. But most protocols are not HTTP.
For that matter, how do you perform more than cursory debugging of HTTP services? Do you seriously sit there and carefully type out HTTP 1.1 compliant requests, along with requisite headers and maybe even cookies? Does that actually work for debugging complex issues, and does it really differ that substantially from the debugging one performs to see if a binary protocol service is up and accepting requests?
1. Debugging which headers were causing Amazon CloudFront to MISS requests originating from Android. tcpdump, observe request, repeat using netcat until the problem was pinpointed.
2. Debugging failed authentication on a dovecot IMAP, on one specific scenario. Again, tcpdump, reproduce, isolate and fix.
When I was a webdev (over the past few years) I did this at least once a month.
I'd daily look at the HTTP headers, though. Also, I'd often add custom headers for debugging purposes.
Edit: It's also not just telnet, it's being able to use simple scripts to automate requests or responses for many reasons.
You can talk about efficiency until you're blue in the face, but human-readability is very useful.
I used to be able to pretty much parse a hexdump of an X.400 P1 PDU by eye (certainly if I could reformat in a text editor), but even today I found it useful to eyeball a recalcitrant programs's HTTP request/response cycle by simply catching it under strace and grepping for "HTTP".
It's a massive reduction of friction to have human-readable protocols. I used to make the argument that it didn't matter and I was wrong.
For who, how often and when exactly? As far as I have understood, we are doing something wrong if we have to deal with debugging protocols which implement an abstraction, rather than just let the abstraction do it's job.
This is like saying that "Well for programmers x86 assembly instruction mnemonics are useful compared to machine code bytes!", to which one could say that for an average programmer that makes no sense.
I didn't quite care about stuff like this until I got introduced to information theory and really started thinking about what it means to send and receive information. The amount of totally unnecessary waste is astounding.
Hi, I'm a sysadmin. Sometimes we have to debug things that are on the other side of the world. When we do that, we need to have a mental model of what we're doing. You're familiar with that as a software developer: you have a mental model of the capabilities of the language that you are using, and you are fitting that in with the mental model of the problem you are solving and the context of existing program code.
To sysadmins, protocols are like programming languages. That's why we like text-based protocols: we can fit them in our heads and type them out, slowly, like ancient creaky teletypes that make lots of mistakes and pause in between commands. Meanwhile, we're running down checklists: can we connect? OK, there's no IP-based packet filtering. What does the banner say? Can it support STARTTLS? OK, disconnect and try again with telnet-tls. Hey, it doesn't negotiate...
Totally unnecessary waste? Sure, assuming someone has already written a tool which does all the testing for you, and you know that tool exists, and you have access to it right now over your tiny smartphone's 3G connection at 0400 while you thought you were on vacation.
As a sysadmin, do you really have to deal with HTTP protocol by hand on such basis that a tool could not do the task for you if the protocol wasn't human readable?
I need to deal with a dozen protocols in an emergency situation where I can't guarantee the appropriate tool is going to be available. (When it's not an emergency, yay tools!)
For the same reason, I can use vi with no vim extensions. For the same reason, I like configuration files written in text formats, not binary blobs. Binary would be faster, sure. When it breaks, you need the precise tools to know what you're doing.
I commend to you RFC 3117 - http://www.rfc-editor.org/in-notes/rfc3117.txt - as a discussion on how to figure out what a network protocol has to do, and how.
Any time you're debugging "one layer above" and you're in Sherlock Holmes territory (i.e. the problem seems impossible, so one of your basic assumptions must be incorrect) you have to check stuff.
And if you check your protocol by using the protocol-parsing tool which comes with your protocol implementation, you're not doing an independent test that it's actually working OK.
As a more concrete example, I'm using a library doing AWS request/responses over HTTP. There are various places I could add instrumentation to dump information but:
- it takes work to add or enable. I can grab the on-the-wire protocol with strace or tcpdump (this a reason why it's always good to provide a non-SSL option for your protocol)
- whatever is causing the problem could be below the layer I'm logging at
- they are all error prone. Maybe I miss a part of what goes on the wire. The data-on-the-wire is the only thing which matters for the protocol. The other end has no additional state.
When you're debugging, you have to validate stuff. And look for patterns. Human-readable protocols facilitate both of those difficult activities, reducing friction.
Instead of ACKNOWLEDGE you say ACK or mere A. Instead of REQUEST PAGE FROM <path>, you say RPF <path> and so on. This is my main point of hatred towards "human readable formats", because they waste bytes for no reason.
I can run the water for as long as I want without absolutely no consequences for me, but why would I do it if I can avoid it? Why would I not save resources whenever I can, even though I don't need to do it? It's more a philosophical question, to which I would answer with "save anything you can, whenever you can and make no waste.". Very simple.
> Why would I not save resources whenever I can, even though I don't need to do it?
It's a cost-benefit. You don't spend your evenings clipping coupons all the coupons you can (which would save you some pennies). Or if you do, you don't stay up late to do it. The benefit to you of some free time is greater than the benefit of saving the pennies.
Similarly, the cost of mild verbosity is balanced against the benefit.
Basically, I don't buy an absolutist position of "save wherever you can". There are costs to saving, make a judgement whether it is worth it.
The problem with many web developers these days is that they don't understand or don't care about the networking side of things. Web development is such a high level view of programming that many developers who've only grown up with targeting the web, those kinds of developers don't also appreciate just how many layers of abstraction there are between them and the users navigating their site. As far as they're concerned, they just bang out some PHP, copy the files onto some shared hosting provider and let the sys admins worry about the rest. Which is fine if that's all they want to do, but there's a whole plethora of technology at work - even beneath the HTTP protocol.
As for tools to query HTTP, I swear by curl:
curl -i --silent example.com | head # http headers (written by web app)
curl -I example.com # http headers (written by web daemon)
curl -v --silent example.com | more # verbose output; great for tracking down faults
curl -H "host:example.com" ip.address # set the host header; useful when using named based virtual hosts
curl -A "opera mobile" example.com # sets the user agent; useful for working around mobile / desktop redirects
...etc. Rarely does a day go by and I'm not using curl.A. base64 everything. Obviously this has high overhead.
B. Escaping (aka byte stuffing). This is somewhat slow to escape and unescape, the overhead is variable (in rare cases 100%), and it's fragile to read or write by hand.
C. Byte counting. This is the most efficient, but extremely inconvenient to write by hand.
You could create an efficient "text" mux protocol (basically BEEP), but the result is so non-human-readable/writable IMO that I don't think people would be any happier. HTTP/2.0 is not complex because it's binary; it's binary because it's necessarily already complex enough that text doesn't save you anything.
What I don't understand is why we need both SPDY and QUIC.
SPDY provides fast multiplexing, compression, and guaranteed SSL. QUIC provides fast multiplexing and guaranteed SSL at the transport(-ish) layer. QUIC does at least one thing that can't be done at the application layer - faster opens - and seems nice to have as a generic base for multiple protocols, so I call it a good thing (and I'd like to see it in the kernel). But the plan seems to be to run SPDY over QUIC - I can't actually find enough information on the Internet (and am too lazy to look through the source) to find out whether packets are going to be double encrypted for the time being, but even if/when that is avoided, the multiplexing and optional encryption seem to be wasteful complexity. I would prefer if HTTP/2.0 were a simple and easy-to-parse protocol providing compression only, expected to be used over QUIC.
So this was entirely by design and probably not all that difficult to implement. That it might not be appealing to some I can understand. I suspect the creators of SGML had very different concerns.
I like that HTML has a very humane interface.
Those people would have figured out how to ship a web page; they'd just would have had so much trouble figuring out why one they shipped was't working, because the browser would have told them.
It's useful for humans talking to each other. Apart from that I agree with you.
Plus, it's SSL based. SSL means you're already used to not looking at things through telnet. Are you really saying we shouldn't use SSL because "it's not just 7-bit plaintext"? That's the point, after all...
Someone has never used `openssl s_client` before.
Very few people actually delve into and edit the TLS protocol. People, including devs, often delve into the HTTP protocol and craft requests by hand, or with a simple text-processing script.
To most people, TLS is transparent. This HTTP 2.0 protocol is not transparent.
Making the protocol I'm working in binary doesn't do me any favors.
As some others have said, I feel that making a "webapps" protocol, a streamlined websocks or something, maybe this HTTP2.0, and keep HTTP for stateless resource representations, like it was designed to do.
The way I see it, we can have our cake, both flavours. Why make someone remember different URLs? "http://" is too entrenched for when people type it in, and I'd rather not worry about browsers trying to sniff out faster protocols as they currently do with SPDY. Better to make it a version number, even if you personally never use it.
(1.0 and 1.1 was all I knew, but Google distracted me in to thinking it was real again with those dang uncorrected news articles...)
And what happened to the calls for SSL everywhere? Defending your cookies from the NSA and MITM? Of course we need two protocols, perhaps more than two. I'd love to see something more interesting happen with multicast given talk of IPTV in 4K.
I'm shocked that your only objection ends up being the protocol name. Worse, it's the version number. But hey, it is semantic versioning... Don't use it if you don't want to. It's not like anyone else will notice if your site is already fast because it's one request with no state.
I dislike both.
> And what happened to the calls for SSL everywhere?
It should be, it's just not HTTPs job to provide that
> I'm shocked that your only objection ends up being the protocol name.
I get upset whenever the next version of something is entirely different in philosophy and design than its predecessors. Use spdy:// or something. You can use the HTTP upgrade mechanism, or something like STS to tell a browser that it accepts the new protocol.
SIP is like HTTP, in fact I believe it was modelled after HTTP and has a lot in common with it. You can troubleshoot SIP issues very easily using a packet inspector like ngrep (using Telnet might be a bit difficult as there are some timers that expire if you don't respond fast enough, but that's besides the point).
Then there's a protocol like SS7 which is binary based. It's all structure binary bit fields. Even though it's binary based, and not very human readable, we have tools that decode the bits into a human readable format, which in turn makes it just as easy as SIP to troubleshoot.
The question I ask myself is, which one would be easier to implement, text based or binary format? I guess I'd lean towards text based--in fact I have implemented a SIP client myself with success. But have yet to implement any kind of SS7 or binary based protocols. Perhaps experience has something to do with it, would having experience implementing binary protocols make it any bit easier? And to be fair, SIP is also much well more documented then SS7. As well it's easier to test SIP.
I'm just starting to work with binary protocols, so it will be interesting how much progress I make. I realize there are also some libraries that help with working with binary/bit field data.
But if you're developing for a protocol with garbage documentation (or no documentation), a text based protocol at least offers some amount of intrinsic documentation.
It's a matter of balance, and we can all choose our own camps along the abstraction <=> ease-of-understanding-every-detail continuum (I'd tend to plop myself toward the former.) Just as we can create a reliable C => assembly compiler, we can create reliable tools to let us leverage abstraction for the better and focus on the bigger picture rather than the smaller.
My "not really" bit is because the better analogy would be C to machine code.
The author's point is that HTTP/1.1 is a text-based protocol. It's inefficient that we have to translate that text into meaningful bits on either end, but it makes it easier for humans to inspect. HTTP/2.0 allows for binary communication, removing the need to translate text, but making it harder for humans to read. It's removing layers of abstraction, not adding them. The introduction is quite clear on this point:
The Hypertext Transfer Protocol (HTTP) is a wildly successful protocol. However, the HTTP/1.1 message format is optimized for implementation simplicity and accessibility, not application performance. As such it has several characteristics that have a negative overall effect on application performance.
I'm not necessarily saying I agree with the author; I don't know enough about the tradeoffs in this domain to say if the loss of abstraction is worth it. But I do think it's worth understanding his arguments on its merits.
My point is that it is sometimes (or from my perspective, often) worth it to deal with more complex (as in higher-level) technologies via tools that allow us to de-abstract those technologies than to focus on using only technologies for which we can understand every detail of their implementation. Trying to understand assembly written by a human vs assembly spat out by a compiler (with its optimizations, etc.) is a very different beast, but thanks to the tools of higher-level programming languages and compilers we rarely need to fiddle around with machine-generated assembly.
Similarly, with reliable tools, we would not need to fiddle around with raw HTTP 2.0 workings, and thus rarely would the concern of immediate human-readability be an issue.
Isn't that directly equivalent to having a preference for a textual wire protocol?
The new vision is the application. Whether it's on a phone, or running as a web site, these are not publishing so much as services.
Perhaps it's time to fork protocols: one for publishing, one for interactive applications.
I feel that the document-centric request-oriented nature of the web has been a powerful influence on the design of "web applications." But we still haven't really figured out how to do it.
The web has been full of applications since the beginning, no? There was no golden age of pure document sharing. So much of the web throughout its history has been CGI, ASP, JSP, ad hoc setups to provide dynamic sites. Session cookies, inscrutable URLs, forms with two dozen hidden fields encoding cryptic parameters...
(Here's a www-talk thread from 1993 discussing the new script support in NCSA httpd http://1997.webhistory.org/www.lists/www-talk.1993q4/0485.ht...)
But since the formulation of the REST architecture, it seems that more and more people are interested in what the WWW was meant to be. Maybe we're only beginning to understand the possible implications of the WWW design, and to purify that into a powerful conceptual framework.
Does that essentially involve a text-based HTTP protocol? I don't think that's the most important part.
The REST philosophy seems to say that the sane platform of the web is more crucially about clearly defined paths to typed resources providing a uniform set of actions and discoverable associations. This is not just for publishing -- what makes it so interesting and fruitful is the way it can be used to structure many kinds of applications in the form of published resources...
"Calling it a day and forking it" seems in some way to mean giving up on this fruitful encounter between two paradigms. Then you get a binary socket protocol with no resource structure but with a lot of potential for shiny stuff, versus a document-centric protocol that's slower and more restricted... And then all the big players do the shiny thing, and the document protocol is left for enthusiastic hobbyists and legacy applications...
True, but many of those dynamic pages actually had semantic meaning in their structure. Think of a discussion forum, for instance. That's a case for which the WWW is perfectly fit.
> Calling it a day and forking it" seems in some way to mean giving up on this fruitful encounter between two paradigms.
On the other hand, this encounter doesn't look too fruitful in applications like Google Docs, Meebo or Prezi.
In the meantime, one side tries to push for things that are primarily relevant for non-semantic applications (think HTTP 2.0's binary protocol, which is all around HN nowadays), pissing off the people who primarily want to transfer semantic content, while the primarily semantic-oriented additions keep dragging the other side back. People have been building web apps with fancy UIs for almost a decade now, and there's still no decent UI builder to speak of.
I wonder what is the best way to do it. convincing all owners of web apps to migrate to a new infrastructure strikes me as extremely hard. convincing the millions of 'publishers' of web writings (including all bloggers) to migrate sounds really hard, too.
There are reasons for all of this, of course -- it's not being proposed on a lark. But it does have a very real complexity cost.
This meme needs to die. It's like ending every post with "Just kidding, I don't know what I'm talking about, I'm just cargo-culting blog attention-desperate person trying to build my Klout score to impress myself."
SOAP? You lucky. It's 2013 and I'm still knee deep in RMI-IIOP sometimes. Wireshark with GIOP dissector helps, but I’ll take SOAP over CORBA any day:-).
Seriously: Your post summarized very well what I was thinking when reading the recent HTTP 2.0 posts. Full Ack.
Think of HTTP 2.0 as a successor for HTTPS instead of HTTP if it makes you feel better.
The output is exactly the same - plain old HTTP.
I don't buy the "argument from visibility", for example.
HTTPS and DNS are both binary protocols and you need tools to parse them in order to get telnet-like visibility. Does that mean when it comes to those protocols "we are forced to rely on other tools to hide the underlying complexity and dumb things down to a point where they are manageable", as the author says?
May I remind these nice people that most text-based protocols we rely on today were designed in a time when CPUs were orders of magnitude slower and modems ran at 2400 baud? Talk about processing and bandwidth resources...
The resources consumed by having to interpret a text-based protocol pale in comparison with the resources consumed by all the levels of abstraction that modern frameworks/language runtimes that we use today. But wait, these save developer time, so they are a good thing. So does having protocols that can be troubleshooted by a human with minimal tools.
"In theory, theory and practice are the same. In practice they are not" :(
https://github.com/nudgepad/space
It's a very understandable and very powerful language that I think will at some point be extended to replace HTTP, amongst other things.
herge: You could also use the "data" program from Hobbit's netcat.
Then he goes off to talk about SOAP, which he hates, because "the only way [he] know[s] to hand-debug a SOAP transaction is with a hammer and a straitjacket." But SOAP is text based, which directly contradicts his argument from earlier.
The other allegation here is that the new protocol "satisf[ies] what are essentially the needs of a small group of very influential players." Again, Marco just states this with no rationale at all. In fact, the opposite is true. By being more efficient, HTTP 2.0 will serve the needs of everyone. Should startups have to pay for 10 servers when they could have paid for 5 with HTTP 2.0? Apparently Marco thinks they should. It's easier for Google to pay for a few more servers than it is for Joe startup. This should be obvious to everyone, but apparently it's not.
The whole binary versus text thing is a complete red herring. If you have ever used wireshark, you know that it makes binary fields easy to read. Binary protocols are easier to develop and standardize. Pretty much every good developer realizes this, and the responses here reflect that. The idea that making something LOOK like (but not actually be) "human text" makes anything better is an idea we should have buried along with COBOL and other mistakes from the past. But I guess not.
REST
http://example.com/resources/item17
SOAP
<?xml version="1.0"?> <soap:Envelope xmlns:soap="http://www.w3.org/2003/05/soap-envelope"> <soap:Header> </soap:Header> <soap:Body> <m:GetStockPrice xmlns:m="http://www.example.org/stock"> <m:StockName>IBM</m:StockName> </m:GetStockPrice> </soap:Body> </soap:Envelope>
Dolphin Mini.
HTC Wildfire.
Android 2.2.2
I guess I'll check it out later.