Why HTTP?
timothyfitz.wordpress.com
timothyfitz.wordpress.com
Also the argument slightly reeks of the same "Everything knows how to deal with it!" that caused some people to use XML for everything regardless of how well it fit the job.
So I agree with the author and my own approach to protocols is to start off with the assumption that I'm using HTTP and only use something else if there's a good reason why HTTP isn't appropriate (performance being the most common dealbreaker).
It's a very specific request/response protocol, which suits some specific purposes.
>> "start off with the assumption that I'm using HTTP"
That's ridiculous. So if you're writing a multiplayer real-time game, you'd assume you're going to use HTTP?
I think he was complaining more along the lines of "why are we running a custom protocol when we're just passing along a bunch of strings?"
I can:
- Write my own protocol with very low overhead, just a simple TCP handshake, the client sends me a little string with its location, and I send back the weather. Simple, really fast, has low network and processing overhead. But then you find that this is almost impossible for someone to integrate into their webapp, or even a native app, without opening sockets and doing a stupid amount of work just to set up that connection. I can also write my server code in C++, which gets me some hellishly fast performance.
or
- I can implement it in PHP, and accessible via a URL like http://www.mysite.com/weather.php?lat=X&long=Y and be done with it. Total coding time? Almost zero. There is no need to set up or tear down connections, and no worry about things like encodings, endianness, or anything of that sort that you'd have to worry about with a raw TCP connection. These tools are also proven, and relatively bug free, which one can't confidently say about anything they re-implement themselves at a low level. Hell, a simple shell script can curl my URL and get the weather back, just like that.
There are gains to be had for doing everything yourself, but for the majority of webapps the tradeoff in reliability and ease of use/development simply isn't worth it. The existence of talent, tools, and community knowledge is often worth the lower performance and sometimes unnecessary overhead.
Sarcasm aside, I don't see exactly what HTTP buys you over establishing a TCP connection to the server and talking to it directly. Points one and two, check. If your favorite language can talk HTTP, it can talk over TCP.
Besides, in a URL like http://foo/bar?v1=x&v2=y&v3=z is in itself a type of protocol, namely the name/value pairs. Sure, I know HTTP and can debug problems at THAT level, but what about at the higher level? There's still a protocol (oh, I forgot that v4 is mandatory when calling bar, but optional when calling baz).
What are we even contemplating building here? HTTP is clearly not a sane choice for everything, and most likely not a sane choice for the majority of things people build.
This argument is like saying "If you're going to write any software, you should assume you'll write it in Java, unless you have a good reason not to". (Which some people were probably chanting in 1998).
Maybe I missed the part in the article where it said "In the world of websites/webapps... etc" :)
I think that the golden rule should be something like:
If you can't do it on http you're free to roll your own but if you can do it over http then you really should.
interesting development:
There seems to be some wisdom here that I am not extracting successfully. It seems HTTP "can" in most cases. But where it can, do you really want to use it? Technically you could also write a db engine that reads over HTTP, but this must be a terrible idea; what would be the argument here that counters the rule of thumb?
- without getting really tortuous, but a little bit rethinking your problem should be fine - without sacrificing a large amount of speed or functionality - without compromising your design goals
If you go to this url you can probably find some contact info to tell the NY Times people how terrible their idea is: http://code.nytimes.com/projects/dbslayer
There are many purposes where HTTP is not suitable at all. Realtime audio/video comes to mind - UDP is commonly used here because even the TCP latency is too much. Anything where you need a persistent, bi-directional connection is not a good match either. Anything where you need to push from server to client. In short: Anything that doesn't fit into the request/response paradigm.
In fact, many consider HTTP in its current incarnation to not even be suitable for the interactivity that we expect of modern webapps.
But ofcourse that doesn't stop the truly enlightened. Hence we got abominations like "Comet", a persistent stream-socket emulated on top of an inherently request/response-based protocol. Or RSS, which is nothing more than Usenet done really, really wrong - all lessons unlearned.
Sorry, got a bit carried away. But you get the idea I guess.
Be careful when you get carried away.
I meant to say: They not only avoid HTTP - they even avoid TCP, too.
Well, the phrasing was not ideal, I'll give you that.
UDP does not guarantee packet delivery, thus HTTP cannot ride on top of UDP.
It would be pretty hard to make http work over something that guarantees transport of a packet but is not connection oriented.
Where Reliable_protocol_wrapper does retransmissions, acks, sequence numbering etc.
There's no reason why you couldn't use HTTP over UDP if you like, I'm sure it's most likely been done in the past - if only just for fun.
It turns out lots of ISPs and broken routers and terrible corporate proxies only allow HTTP. They do things like require HTTP headers, so they can check for whitelists/blacklists. It's stupid and it breaks TCP/IP in general, but most people only care about the internet so it just works for most people.
To me comet is a testament to our collective failure to demand proper tools from the "big guys" (Mozilla, Opera, Microsoft). The term "web 2.0" was coined roughly 7 years ago and even long before that it has been clear that some sort of WebSocket would be incredibly useful to have in a browser.
Yet no single browser vendor (to my knowledge) has impemented it today.
Imagine where we could stand if only Firefox had a true, non-standard Socket class already. People would be using it (socket for the fox, comet-hacks for the rest) and IE market share would probably be dropping faster than ever because all the fancy new apps work better elsewhere.
But I digress, this is mostly whining. You have a point; the comet hack, as ugly as it is, is justified by the lack of proper alternatives. But still, an abomination on so many levels...
You say "all lessons unlearned" implying that RSS is worse than previous incarnations of solving the same use case. This is a common opinion I really am railing against. RSS is EXTREMELY popular. Usenet is a dying architecture. It's a case where being worse in the "obvious" technical architecture allows you to be better in things that actually matter: ease of access, simplicity of use, lack of installing things and wide support. These are just some of the reasons RSS took off. Easier means more viral, because easier means a shorter viral loop.
Of course, I think RSS is better for syndication than notification, and would prefer more sites like twitter use Webhooks in addition to RSS.
Well, ubiquity != adequacy. I understood the original author's article as a technical recommendation. He basically suggests to use HTTP for everything because he thinks "it's technically good for everything".
He made broad claims about how any HTTP based protocol will magically scale by leveraging proxies, loadbalancers and other existing infrastructure, completely ignoring the fact that many applications just don't fit into the request/response paradigm in first place.
RSS is a great example for the power of the internet that enables us to "build on what we have" without waiting for some standards body or greater authority to get moving. But it is also a text-book counterexample to his scalability and "one size fits all" claims. Polling just doesn't scale for these things and technically it's a step back from Usenet, that had these problems sorted out already. We went back to square 1 with RSS and are now locked up there until we get a true WebSocket and worthwhile persistent storage in browsers.
As an example where HTTP doesn't work so well, is pushing information as opposed to servicing requests. COMET is a hack and requires you to be very careful with clients. (e.g. the 2 connection limit in older browsers) You'd only really want to use it where there's no alternative, like in a browser.
Basically, HTTP is great for servicing individual requests. It can do kind of do persistent connections, but they're more an optimisation than part of the design and therefore optional for the client even if the server suggests them.
If your use case falls outside of the basic request/response premise, you can maybe use it for prototyping for a while, but you should probably switch to something more suitable to your setup after that unless you need to use HTTP for reasons outside of your control.
Sensor networks are a bunch of sensors that collect info, then network themselves to transfer what they know, usually to a mothership base station of some sort. Because it's a pain in the ass go out in the field and replace the batteries on these sensors, you have to make their transmissions energy-efficient. There is nothing in the HTTP protocol that takes energy consumption into account. Nor should it. It's not an application layer problem. However, the only reason I mention it, is that I've seen people start mixing the protocol layers together in a single protocol to achieve the energy consumption characteristics they want, which included the application layer.
So hence, that's one instance where HTTP isn't a first choice. As long as you know what HTTP is for, and don't use it blindly, you'll find that it probably suits your needs.
We choose to implement a CPU efficient binary protocol, automatic client-side queuing of outbound messages if a server failed, and automatic -client-side- fail-over to the next available server.
The initial implementation, with no time spent on profiling/optimization, was able to receive and process 25,000 event messages per second from a single client, and scale up clients/cores relatively linearly.
I can't even begin to fathom solving this problem with HTTP, or why the 'features' of HTTP listed (proxies, load balancers, web browsers, 'extensive hardware', etc) would be an improvement over the relatively simple and highly scalable single-tier implementation we created.
Clearly, HTTP works well for some things, but "just use HTTP, your life will be simpler" is a naive axiom.