Unless the protocol specifically guarantees certain behavior and commonly-used systems regularly exercise this guarantee, it's just not going to work reliably when it's needed.
Hearing some of the "works for me" discussion from developers suggests that we're heading for that magic situation where it works 99.9% of the time. I.e., the system looks fine in testing and then fails in mysterious ways (that require deep protocol fixes) in production.
Ideally, implementations of such a protocol would intentionally fragment the messages somewhat if they were not going to guarantee they were atomic. But there are very few developers (and code reviewing managers) enlightened enough to let that kind of thing ship.
Even in cases where your WebSocket implementation exposes incomplete message fragments, the boundaries of the message are still clearly and accurately preserved.
Luckily there's a rather excellent compliance test suite, here: http://www.tavendo.de/autobahn/testsuite.html which should go a long way to help nail interop issues.
So that's a design tradeoff. I've implemented protocols that did it both ways and it's definitely easier on the guy trying to implement a library or other receiving application if he can get a reasonable upper limit on the size of the messages.
But on the other hand, by not requiring the total message length be known in advance, it eases the logical (and memory) burden on the sender. Often the sender will be an overloaded server.
Nothing can prevent a higher-level protocol on top of WS from negotiating its own max message size.
So the design choice that was made would seem to allow optimizations for the overloaded server case without prohibiting other optimizations. This is typical of W3C protocols.
So, if I can read 2 bytes from my buffer, I can get a good idea of the message size. If I only have 1 byte, I rewind and wait until I receive another.
If Websocket developers begin making unwarranted assumptions about message framing and fragmentation you will regret it. BELIEVE ME.
/me gets back to bugfixing
There's two ways to respond to this: 1) claim that websockets is broken and all client APIs should be streaming or 2) not unexpectedly send multi-gigabyte messages.