WebSockets is a stream, not a message based protocol
lenholgate.com
lenholgate.com
Time and time again it has been demonstrated that we are bad at choosing a maximum allowed value for all applications and all future considerations (see: ethernet frame sizes, IP address lengths, operating system address spaces, file system block sizes/counts, etc).
In some cases (many of those previously listed) there were hardware, cost, or technical concerns that led to nailing down a number in an RFC. For WebSocket there is no clear benefit to forever encoding a specific numeric maximum message size. It is a high enough level protocol that there is no technical or cost benefit to make message sizes limited by anything other than individual application needs.
As such, the WebSocket RFC leaves maximum message size implementation defined, and specifically says that an implementation SHOULD implement a reasonable maximum message size for its purpose. A chat application that knows it will only be moving small text messages can set its maximum message threshold small to improve buffer performance and catch invalid messages sooner. An application that finds a business case for sending a large file in one large message can set itself up accordingly. Generic WebSocket parsers should expose a method of setting the maximum message size the application wishes to receive.
I definitely agree that not requiring implementations to return their maximum message size along with the "Message too big" error will make some sorts of interoperability more difficult. However, it also prevents exposing implementation security details and simplifies the core spec (the author has already complained that the spec is too complicated already). It is relatively simple for an application to negotiate a maximum message size privately if necessary and the WebSocket extension mechanism allows a method for standardizing a way of doing so if this turns out to be a serious issue in the future.
The draft in question suggested that providing a message based interface to application code was possible and that the parser could/should deliver only complete messages to the application code. That's hard to do if you also want to allow for the 'endless streaming' scenario that others on the working group were fond of. The result was a bit of a mess.
The final RFC addresses some of this, but there's no getting around the fact that the websocket protocol itself can't tell you how big a message is until you get the final frame.
Sure you can work around all of this even for a generic parser but the initial wording in the draft in question could lead you towards the wrong design if you're not careful.
I do see your point on the "endless streaming" section of the RFC. Stating that "(section 5.4) The primary purpose of fragmentation is to allow sending a message that is of unknown size when the message is started without having to buffer that message." implies that a web socket implementation should support this sort of operation. Indeed, if you want to support sending messages of unknown size you must expose an interface more complicated than the default message based one.
That said, a message only implementation that does not allow sending unknown sized messages is 100% compliant with both the spec and receiving such messages. The RFC probably could have made this fact more clear. I believe that endless streaming mode will not be a common use case and have not implemented it in my generic WebSocket library. I do believe, however, that fragmentation of messages provides important benefits even without unknown size sends. Once you have message fragmentation there is no additional protocol cost to allow unknown size sends.
The wording of the RFC has improved since that draft and the flexibility could be useful in some scenarios.
I ended up with an API which can be asked to deliver complete messages 'if possible' given the buffers provided by the client of the API. If it's not possible and the buffer becomes full then the API simply gives you the fragment of data and tells you if it knows how much more there is to come or not.
The problem is that the original handshake wasn't good enough (there were still security vulnerabilities despite he handshake), and when Ian Hickson decided to hand over control to the IETF, the architecture astronauts took over, adding complex framing with six different frame types, subprotocols, extensions, versions, complex bit twiddling required to parse frame headers, fragmentation of messages into smaller frames (which is what this article is complaining about), control frames interleaved with fragmented messages, numeric status codes and textual close reason strings that "MUST NOT" be shown to the user, masking of data by xor'ing with a random value that changes for each frame, but only for one direction (client->server), a two-way closing handshake on top the existing TCP mechanisms for closing the connection, pings to test the connection for liveness, and so on. There are six registries defined for IANA to keep track of http://www.iana.org/assignments/websocket/websocket.xml; extensions, subprotocols, version numbers, close codes, opcodes, and framing bits.
And despite all of this over-engineering and attempt at extensibility, all extensions must know about each other, because there is no standard method for delimiting different extensions' data (or even specifying how much data an extension uses), and there are three header bits and 10 frame types that all extensions must share. And I don't really know why there's a need for subprotocols on top of the ability to just encode that information in the URL.
It's kind of sad how what could have been a relatively simple and easy to implement protocol has been taken over by architecture astronauts. Yes, a few of these features are actually required to securely deploy websockets (the handshake and masking). Most of them are people making up features that would be nice in theory, instead of implementing something simple that works. Ian Hickson's original protocol wasn't perfect; it still needed some work by the time he left. But it was simple, and easy to implement, and didn't impose restrictions that couldn't be worked around at a higher level.
The really nice thing about SSE is that you can fall back to long polling very easily with exactly the same back end and as it runs over vanilla http without the upgrade protocol system is much easer to implement, you just don't close the connection after sending a message. Obviously its only one way but we have a well established way of sending messages in the other direction with http POST.
http://www.caniuse.com/#search=eventsource
However, since SSE is HTTP-compatible you can easily implement fallback for these, e.g.
Thanks for the info, everyone.
Firefox, Chrome, Safari and Opera all have it. I believe IE10 will have it but I cant find a reference right now.
This seems to be the nicest shim to let you use EventSource on older browsers https://github.com/Yaffle/EventSource, I did make something similar myself but this is much more complete.
For example a collaborative project management tool like Asana where saves are silent and update live on other peoples screens is a better place for SSE and a REST api than WebSockets, where as a real time game is probably a good place to the latter.
I just launched the beta of a small service for making Server Sent Events painless at http://eventsourcehq.com
It's currently only available for members of Heroku's beta program, but I plan on launching it as a stand alone service as well. In any case all the code behind it available on Github...
Everything you send up or down is a message, or a packet, and the size of that cannot exceed the size of pipe (with bottlenecks or intermediary restrictions). Call 'em Quantum Packets, or a stream, or a message. The websocket protocol, as imagined by this developer, is meant to allow a continuous lot of Quantum Packs to "flow", without the application level overhead of parsing a bunch of protocol, headers, wackness. I want to get the data into my applications AFAP, cuz I still have to transcode it, analyze it, and all else to make the baby dance.
What we need as developers are minimum-for-reliability standards. No two people in different locations will have the same pipe. As a developer, I consider it my domain to write software on top of, or using, the socket layer to determine the potential through-put of the given socket, and to test such as needed through-out the simulcast. I don't even want the socket-layer-wrapper writers (may God shower them with blessings) intervening at this level, until everybody on Earth has unrestricted 10mbps/s up and down.
If that control is hidden from me, or not an option, or is nullified by protocol, then my app or media could break in ways I could not predict or understand, and so I would have to design my app using the socket layer in a lowest-common-reliability kind of way.
These are not the opinions of a WebSocket RFC acquainted developer.
Unless the protocol specifically guarantees certain behavior and commonly-used systems regularly exercise this guarantee, it's just not going to work reliably when it's needed.
Hearing some of the "works for me" discussion from developers suggests that we're heading for that magic situation where it works 99.9% of the time. I.e., the system looks fine in testing and then fails in mysterious ways (that require deep protocol fixes) in production.
Ideally, implementations of such a protocol would intentionally fragment the messages somewhat if they were not going to guarantee they were atomic. But there are very few developers (and code reviewing managers) enlightened enough to let that kind of thing ship.
Even in cases where your WebSocket implementation exposes incomplete message fragments, the boundaries of the message are still clearly and accurately preserved.
Luckily there's a rather excellent compliance test suite, here: http://www.tavendo.de/autobahn/testsuite.html which should go a long way to help nail interop issues.
So that's a design tradeoff. I've implemented protocols that did it both ways and it's definitely easier on the guy trying to implement a library or other receiving application if he can get a reasonable upper limit on the size of the messages.
But on the other hand, by not requiring the total message length be known in advance, it eases the logical (and memory) burden on the sender. Often the sender will be an overloaded server.
Nothing can prevent a higher-level protocol on top of WS from negotiating its own max message size.
So the design choice that was made would seem to allow optimizations for the overloaded server case without prohibiting other optimizations. This is typical of W3C protocols.
So, if I can read 2 bytes from my buffer, I can get a good idea of the message size. If I only have 1 byte, I rewind and wait until I receive another.
If Websocket developers begin making unwarranted assumptions about message framing and fragmentation you will regret it. BELIEVE ME.
/me gets back to bugfixing
There's two ways to respond to this: 1) claim that websockets is broken and all client APIs should be streaming or 2) not unexpectedly send multi-gigabyte messages.
My impression of WebSockets is that it's not actually a "finished" high level protocol. They could have just brought a basic socket style interface into JavaScript and left it at that. (And based on its name, that's what you'd expect at first.) But they decided to add various features, (for better or worse, I don't know yet) on top of that. (I guess part of it is the challenge of working not just on TCP, but sort of within HTTP as well). Just as you wouldn't just pick up TCP and start blowing "data" through it without some additional application specific structure, you're going to need to add your own structure inside WebSocket's framework.
The wording was improved around the suggestion to provide only a message based API.
I think the WebSockets protocol ended up being a little more than it should have been. You have to understand that it was being pulled in all sorts of directions by the working group members and that there are good reasons for all of the parts of the protocol (though some of those parts could work better with other parts IMHO). It had to be finished at some point though and I think the working group did a good job in the end.
Personally I think it would have been better had it been explicitly stream based from a user's perspective, but then I don't have the javascript/browser background to know how foolish that probably sounds.
Personally I wouldn't have included the 63 bit message size.
Why you'd want to do that is another question entirely. Introducing roundtrips by feeding tiny chunks to TCP is generally a horrible idea, however, it does prevent the server from dedicating a potentially huge chunk of RAM to buffer the result ahead of time.
Because of this feature, and the author's desire to model this feature as part of some client library API (a mistake? you decide), he's concluded that it's in fact a stream-oriented protocol. That's like concluding it's a byte-oriented protocol because TCP can/will further fragment the partial frames due to segment size constraints, etc. (i.e. it's a silly conclusion).
As I said, the draft at the time suggests presenting whole messages to the application layer. The parser can't know it has a whole message until it gets the final frame... This could lead to interesting memory usage ;)
The protocol provides for a series of infinitely long messages, each separated by a terminator. I don't have a problem with that in itself, but the draft at the time was misleading to suggest otherwise...
If, instead, you just explicitly say that it's a message oriented protocol, then the software that implements it (both on the client and server side) can just provide an API that delivers a message at a time, and if anything happens to be fragmented, they deal with buffering and reassembling it, rather than depending on the application author to get that right.