A Badass Way to Connect Programs Together
joearms.github.io
joearms.github.io
> I think it is totally crazy to send JSON or XML “over the air” since this will degrade the performance of the applications giving a bad user experience and higher bills - since ultimately we pay for every bit of data.
If every JSON and XML API switched to OSC, the difference in data usage would scarcely be measurable. First: The vast majority of mobile data is video, audio, images, and application updates. Second: When sent over the wire, JSON and XML are typically compressed with gzip or deflate.
And let's not forget: JSON and XML have some significant advantages over most other serialization formats. They're ubiquitous. Practically every programming language has libraries to parse and generate them. Practically every programmer knows how to work with them. They're human-readable. No special programs are needed to view their contents.
In other words, JSON and XML are only "wasteful" if you don't count developer time. Once you do, it suddenly makes a lot of sense to "waste" hardware resources. After all, programmers are expensive. Hardware is cheap.
Except in China. I worked at a US company that did almost zero optimization on their code, deploying to 5000 AWS instances with a team of about six. In China, bandwidth closely followed by CPU were the cost centers, and the Chinese team of 30 engineers optimized the same code to run on around 50 (admittedly beefy) servers.
But that said, it's rather a good thing to point out for the OP, as Joe Armstrong himself tends to value programmer time over a machine's, and that micro-optimizations are rarely ever worth it (as opposed to yours, which sounds like there were huge, glaring algorithmic issues to fix with macro optimizations, complete replacements of algorithms rather than tightening up code, careful profiling and tweaking to get a 30% speedup in one critical section, etc)
However, getting X hours of average programmer time, vs X hours of average machine time...machines are cheap. You should default to optimizing programmer time, not machine time.
I just moved a customer off EC2. A small startup. I've billed them about $20k for the work. It will take them ~2.5 months to repay in saved hosting costs at current load levels. But if they're still at current levels in 3 months time, something is wrong. On top of that their ops costs have dropped as we have more control over the environment.
Basically, with rapidly growing hosting needs, getting a more cost effective setup was a matter of survival: They'd be unlikely to close another round of funding in the next few months if they didn't get that cost under control.
I wonder how many startups fail because they don't understand how to get their hosting costs under control, as it's way too common that I see developers that seem to think that servers are basically free, and managers that have no clue they need to seriously question why developers are making the server choices they are making..
But, as we see with Twitter, it often is completely ignored.
Luckily WhatsApp gave some inspiration for companies to reverse this ideal.
Obviously, 4950 units of hardware, repeated monthly into perpetuity, compared with 30 units of developer, for a fixed length project, is not cheaper.
Perhaps he goes a little bit far, but I'd agree that shoehorning your data serialization problems into a ubiquitous format can be a real headache.
One of the interesting things about Erlang is it has binary pattern matching. Matching any binary in it is just as trivial or even more trivial than parsing json in other langauges.
Here is what an IP packet parsing might look like:
<<4:4, HeaderLength:4, _DiffServ:8, _Length:16, _Identification:16,
_Flags:3, _FragOffset:13, _TTL:8, Protocol:8, _HeaderChecksum:16,
SrcAddr:32/bits, DstAddr:32/bits, OptionsAndData/bytes>>
It looks almost like casting a binary blob to a C struct, but you can do even crazier things like take Length and match it later in the pattern, if it specifies a length prefixed binary, for example.> the difference in data usage would scarcely be measurable
It depends. See, Joe's background is from Ericsson. Their stuff controls access of 50% of world's smartphone <-> internet traffic. When people say "but who uses functional languages in the industry?" the answer is if functional languages won't work, chances are high, you won't see picture of cats on your smartphone.
But the idea is that mobile bandwith is still a precious resourse. Even if our CPU and hard drives are getting bigger. That is one shared resource that is expensive. Maybe if a 1M clients use JSON vs a binary protocol doesn't make a difference much for some startup, when we talk about billions of connected devices talking over radio waves, those things add up. Well, also, game servers use UDP and binary often because of latency. Facebook knows a thing or two about chatting with lots of users, so they use Flatbuffers https://code.facebook.com/posts/872547912839369/improving-fa... that's binary as well.
I wrote an i3status replacement that does exactly this. It sends JSON to STDOUT at a set time interval. i3bar takes the JSON as it comes through STDIN and changes the look of the bar as it's read.
Again, the vast majority of mobile data isn't JSON or XML. If mobile bandwidth was a Californian drought, video would be alfalfa farming. JSON would be taking a long shower: maybe you feel a little bad, but it doesn't matter. All the short showers in the world can't make up for the huge volume of water (bandwidth) used by alfalfa (video).
Obviously, if you have a billion users, trade-offs change and it can make sense to use esoteric or bespoke protocols. But even for Facebook, using JSON isn't crazy. It's just inefficient. And Facebook didn't switch to FlatBuffers for the bandwidth savings. Their primary goal was to improve loading times for their local cache. Had FlatBuffers required 10% more network traffic than JSON, they likely still would have made the switch.
Sorry. It was mainly a response to "JSON and XML are only 'wasteful' if you don't count developer time. Once you do, it suddenly makes a lot of sense to "waste" hardware resources. After all, programmers are expensive. Hardware is cheap."
And pointing out that it really depends on what languages or type or programmers you use and what "hardware" is. In some languages binary parsing is really easy and just as cheap and elegant as using a json parser. And well, Joe is Erlang's father, so point out Erlang was obvious choice.
As for hardware, bandwidth is not hardware. It really is "something else" -- a shared resource usually. And if there are enough devices using this even small changes start to make a difference.
Think a bit like a nested loop -- a small optimization in the inner loop might have large visible benefits. Or say, a huge database with 10B records, choosing an 32bit int vs an 16bit one for a column will make a for a good difference in size. That is basically where Joe was coming from. Really valuing bandiwidth as a shared resource. (He also likes to talk about protocols as well quite a bit).
> Again, the vast majority of mobile data isn't JSON or XML
The important data is JSON though, the control message, the pages that load (and they are slow), stuff that stops, start, manages everything all that is JSON.
> If mobile bandwidth was a Californian drought, video would be alfalfa farming
Hmm, video is 50% of all traffic on smart phones. (http://www.ericsson.com/res/docs/2015/ericsson-mobility-repo... says 45% in 2014). Not that the rest is all JSON but, I wouldn't say it is all just "taking showers" type of equivalency, it is more like grapes, almonds, strawberries, and walnuts.
Ignoring everything else there are much lower hanging fruit to save bandwidth.
You are right that for the majority of applications json is fine.
Isn't that the whole point of compression: to minimise repetition of common features in data?
ASCII is extremely repetitive. 1MB of random braces gzips down to 163KB. That's a compression ratio of 6.1:1. Optimal compression would be close to 8:1 (since each byte really contains only one bit of information).
Your sample data is bunk for determining the compression characteristics of braces in a realistic json file.
Huffman coding works by replacing common symbols with shorter codes and uncommon symbols with longer codes.
The codes used are therefore variable-length.
The codes are created in such a way that no code is the prefix of any other code.
That way, the decoder is able to know when it has reached the end of a code without the need for extra information other than the Huffman tree, which tells the decoder which codes belong to which symbols.
The Huffman tree can be pre-agreed or, more commonly, included with the compressed data.
JSON/XML are the "dynamic language" of the transport protocols: flexible and quick to prototype, but can become very hard to "refactor" and performance suffers, as the thing grows.
As for public APIs, obviously agree with your point - it's a lot easier to roll out a JSON interface as parsing libraries are widespread, etc etc
The difference between communicating changes by setting a bunch of rules on how to do versioning/structure/etc vs changing an IDL is akin to verifying a program structure with unit tests and code reviews vs enforcing them with a compiler (sure, you can always use stuff like Json Schema or XSLT - personally never had a good experience w/ those). The latter is far easier, in my experience.
I've worked in places where binary serialization + deserialization functions were generated automatically by just adding a field and the bit-mapping into a message-spreadsheet. This was way more productive than misspelling the property name of a json-field between serializing and de-serializing and spending a whole day troubleshooting that. This also automatically documented the format instead of sending an example xml-file to your counterpart who is to implement loading of the data elsewhere and saying "the data will look aproximately like this". Viewing the data for debugging was done with a special program (similar to wireshark) that understood the protocol and rendered a much nicer view of the data than pasting a json-file into notepad++, there you need indentation and highlighting plugins anyway so the number of programs you need installed is not an argument either, streaming data through notepad++ doesn't work so well either if you want to view data on the fly.
Security is also one aspect, at every place i've worked there is always one smartass that thinks xml can be created by string-concating, that's impossible if the serialization-functions are totally encapsulated.
All this isn't a problem inherent to xml/json as it can all be solved with proper processes, it's just that it invites to that type of culture.
…which probably amounts to CPU-years of wasted time and electricity, considering how ubiquitous this is.
There are probably lots of things we do as programmers that are insignificant in the context of our own systems, but are so culturally ubiquitous that they add up to an enormous impact globally. I wonder how even things like different application architecture would consume different amounts of power. Or,is there a side effect of shifting computation to the client side of causing more energy to be consumed from dirtier sources, if, say, data centers tend to draw from cheap and clean hydro power.
Sending a jpeg, over compressed http(s), throught a compressed and encrypted VPN is too much compression. Then the client may save the jpeg to a local compressed NTFS cache... This is just wasted performance, plus compressors are a common place for bugs, so this is also a security issue.
Just stop adding compression to everything and think about it first, and then Mobile phones and modem lines may very well be an exception.
Also there is a theoretical knowledge that nothing could beat messages of fixed-length headers with adjustable-length payloads of binary tagged data.
Made of: 1) a suite of interconnectable videogames & apps and 2) physical patchbays to connect them
Came from thinking about software like modules in a modular synth. Playgrounds of interconnectable control structures. Was patching games into samplers, slaving text editors to drum machines, etc.
related writeups: http://www.illucia.com/faq/ another from 2010 or so when I was just starting coding (so a lot of it feels silly now.. but still some related nuggets): http://www.paperkettle.com/codebending/
That said it was great fun building a robot with the controller but hard to get asynchronous feedback. If I were doing it again I'd try to figure out a full duplex OSC channel for something like that.
Oh so true. I wish I learned Erlang (and to a lesser extent Go) much earlier in my career after wasting so much time with Node. I now laugh when I see people on Twitter/Reddit talk about Node's concurrency as a reason to learn it over other languages.
The problem with Node.js and concurrency is that everything depends on trusting code to be perfectly written, and that perfectly written code must be written in a naturally confusing callback style that is really easy to screw up.
If you do it perfectly, then you get great scalability. But one bonehead mistake will ruin your concurrency. By contrast with a pre-emptive model you get decent scalability really easily, but now you've got a million and one possible race conditions that are hard to reason about.
This is not a new design tradeoff. Go back to the days of Windows 3.1 or the old MacOS versions below OS X. They all used cooperative multi-tasking, just like Node.js. Today what do we have? Pre-emptive multi-tasking in Windows, OS X, *nix and iOS.
Web development has actually gone back to a model that operating systems abandoned long ago. As long as your app is small, there is a chance that it will work. But as your app grows? Good luck! (That is why operating systems uniformly wound up choosing hard to debug race conditions over predictably impossible to solve latency problems from cooperative multi-tasking.)
You also don't get shared memory concurrency when it's beneficial, and you have to speak in callbacks.
There are better langs out there for concurrent web. Erlang and Haskell+Warp are fantastic, for instance.
Think proxies (haproxy), routers, forwarders, web servers (nginx) etc. Where memory context per connection should be minimal, and everything should be as close as possible to the select/poll/epoll loop.
Funny enough, this also includes demos, quick example scripts, and benchmarks. I wonder if that what hooked most people to the reactor pattern -- small examples in Twisted, Node, etc, look pretty easy and simple. But when you start adding business logic and callback chains evolve into callback/errback trees 10 levels deeps when things get very scary.
Right. Yeah there was a presentation about Java concurrency patterns about that. Basically dispelling the myth that is everything non-blocking and asynchronous will be faster than the old school blocking thread / socket.
Interestingly, I believe, haproxy and nginx use a hybrid model. They have a worker thread per CPU, all listening to the same socket using epoll from multiple threads! When data arrives they all get woken up and then use a shared mutex to decide which one will handle the request. (Now later kernels fixed that one can have exclusive wake-up across the same socket).
There is a so-called accept_mutex, but I'm pretty sure you can't avoid that if you want to have multiple cores handle connections from the same port. Even Erlang would have to do that somewhere behind the scenes. Newer Linux kernel versions support the SO_REUSEPORT which is meant to address this situation - I guess this is what you're referring as exclusive wake-up accross the same socket?
The problem with synchronous I/O is not speed, but blocking. Blocking means the only possible concurrency model is based on processes or threads. Consequentially, most performance problems with synchronous I/O stem from the thread model: context switching, synchronization, memory overhead, thread-pool starvation and so on.
...all stuff that people should not implement in javascript in the first place.
Co-routine context and and green thread stack size can be tweaked and optimized sometimes, but if you want to have precise control over memory allocated per-connection, reactor/proactor is hard to beat.
Besides, to be 100% honest, C and C++ just doesn't have native support for any other concurrency model. That's probably the main reason why Node.js was designed to use callbacks. I'm pretty sure it would use generators or async functions if it was designed today.
Basically almost none of the criticism in this whole thread is true today. The only remaining true bit is that node still relies on cooperative multitasking and one of the tasks can hog the CPU of a single worker. Which isn't very good, but still, way better than say, the good old Rails 1 request per process model.
Someone should probably do a proper benchmark to show how different numbers of workers at different CPU workloads affect a node service's response time / latency. Especially with multiple processes (cluster), I would bet the effect would be much better than what people expect it to be.
"Practicing" different concurrency patterns sounds like a lot of tedious effort, and the suggestion to "go back to Node" seems a bit high-handed to me. I would encourage everyone to pick and choose what expertise they want to develop.
The pattern isn't special or strongly correlated to node.js
That predates the open sourcing of Erlang by a few years.
Also ACE didn't originate in SV.
A usable reactor pattern requires green threads (or threads) . Node.js does classic IO multiplexing (i.e. what has been available in C since the introduction of select()). It's not bad, but don't delude yourself into believing node.js does anything new.
As a non-fan of Javascript as she is wrote I'd hear this criticism and nod in almost any situation, but not while the author is simultaneously advocating writing your stack as a conglomerate of processes written in different languages communicating over UDP.
[1] https://developers.google.com/protocol-buffers/docs/encoding
A leading thought in a lot of this is the "unset and a zero value are equivalent". It simplifies a lot of what the serializer and deserializers have to understand, and places the burden on code generation or libraries rather than complex serialization. It was unnatural for me to think this way, as I take the opposite for memory (record whether something is set or not, don't just assume it equals zero), but on the wire this makes for very compact and backwards compatible messaging.
Cap'n Proto[1] I have not played with, but would love to learn more about to see the advantages and disadvantages between protobuf, msgpack, and capn.
[0]: http://msgpack.org/
Message Pack has a specification that is well documented[0] and even the rfc you refer to specifically states it has different goals (though, eliding specifics with pretty sad generalizations)[1]
You find those generalizations sad because that is the acknowledgements section!
If you start from the beginning, you only need to read three paragraphs to see that "Appendix E lists some existing binary formats and discusses how well they do or do not fit the design objectives of the Concise Binary Object Representation (CBOR)." [0]
If we look at Appendix E, [1] we find that section E.2 contains an explanation that you're likely to be more satisfied with. [2] However, before you go off and read section E.2, I strongly urge you to read the couple of paragraphs in Appendix E, first. The authors of RFCs generally tend to try hard to remove redundancy in their prose, and later sections often elide information covered in earlier sections.
[0] https://tools.ietf.org/html/rfc7049#section-1
Things like stating that "evolution has stalled" while also recognizing that the format is stable is hand wavy when you consider we're discussing things that go over the wire and even end up on disk. Yes, stability should be a goal.
The real difference between CBOR and MessagePack is that CBOR wants to be "schemaless" in the applications themselves instead of just on the wire. They hold up json as the example format for something that doesn't require schemas, and yet I see "json schemas" being published[0], and even people trying to standardize the schema format[1]! Looking at any modern JSON API would tell you that "schemas" in the xml sense are not required, but applications all must be very knowledgeable of the format.
Having a data type for "PCRE" is just insanity on the wire, and I can't imagine the type of API you'd be publishing where you would accept URLs or Regular Expressions or Text or Binary, AND want to be able to decode them into proper types in memory all without applications on both ends knowing that ahead of time.
Which brings me back to my initial point: CBOR is not just a "standardized Message Pack", it's a very different approach to what they think the applications on either end of a protocol should be doing.
[1]: tools.ietf.org/html/draft-zyp-json-schema-03
Protobufs depend, for interpretability, on a schema. That is, if you don't have the schema file (which is normally communicated out-of-band, e.g. by source control), you know almost nothing about what the wire-encoding should decode to.
Protobufs use variable-length integers (varints), both as a datatype for the payload, and as an essential feature in understanding the encoding.
The wire keys in protobufs don't fully specify the datatype. Instead they include 3 bits to tell you what length-scheme is used, and some other bits (variable number of bits, because varint) to tell you which field in the schema will tell you the rest of the datatype for this field (as well as the key-name of the field).
Protobuf is a rich protocol for doing potentially-complicated things. This OSC thing is very nearly the dumbest thing that could work. Indeed, Armstrong's post makes a point of noting that. To him, that's a feature.
I agree this might not be the best method of dealing with out of order reception of packages, but it does exist in the protocol.
I think the point is just to say that the article shouldn't tout plain UDP as a fix for TCP being hard in certain languages without calling out the downside.
That said, most protocol implementations that I'm aware of over UDP have one or more of "retry/order checking/etc" on top of them. Like NFS, Bootp, tftp, etc.
But it's a common false belief that it would be "reimplementing TCP". Some guarantees of reliability are required for most practical applications but TCP's "in-order stream of bytes" isn't suitable for everything.
For any kind of real time data, when packet loss occurs, TCP will re-send data that may already be outdated, and causing later packets to arrive even later. In real time uses, TCP makes any networking problems worse.
It's also worth noting that under ideal conditions (ie. no packet loss), TCP and UDP behave almost identically (after startup, that is). The issues only appear when you're working in less than ideal conditions and are worse the longer the physical distances are.
You shouldn't drop packets on single machine unless you are sending a lot of them and run out of buffer. Buffer size is also tunable if you do have issues like that.
If you need to do things over the internet it would probably be best to do OSC over TCP.
[0] https://docs.google.com/document/d/1lmL9EF6qKrk7gbazY8bIdvq3...
Nice to see a good write up on this project!!
http://en.flossmanuals.net/pure-data/ch065_osc/
You may be interested in qmail and its sub-applications, which does it properly IMHO. It's a great example of the article's key point, which is the value of simple messaging.
However, I was thinking of it more from a Amiga COPPER instruction list perspective, but it amounts to the same thing.
The really cool thing is that approaches like this tend to be language-agnostic, easy to program to, and extremely fast. All points raised in the article.
And while it's request-response based over message ports, personally I see AREXX ports in most larger Amiga applications as a good example of the overall approach of cooperating processes to build larger systems.
Dbus and similar systems today has lots of great functionality, but what they miss is what ensured AREXX was "everywhere": It was trivially simple to support.
Pointing to the simplicity of OSC in the article really resonated with me, because a lot of these systems are far less useful than they could be because it's extra work. Simple systems on the other hand end up shaping application structure:
AmigaOS was massively message based from the outset: Message ports and messages were "built in" and used all over the place. E.g. when communicating with filesystems or drivers or the GUI you do it via messages.
Then came AREXX and there was suddenly a standard message format for RPC - both application-to-application and user-to-application.
If Amiga-applications were message oriented before, a lot of "post-AREXX" Amiga applications took it to the next level by being shaped around a dispatch loop that included not only e.g. GUI updates, but defining internal APIs that exposed pretty much every higher level action in the application, and used that as an explicit implementation strategy.
It is really fun to use these beautiful Star Trek looking interfaces to make your robots move around for example ;)
Seems like a recipe for inevitable incompatibilities when two languages no longer have codecs for the new underspecified layered format.
The layout of controls is hierarchical in the protocol; each control has an associated endpoint, which looks very much like a unix file path. For example, to unpause an audio player, you would send the message "/player/is_playing/set t" (or f, to pause). If you wanted to control the bands of a 10-band equalizer, you could implement that as "/player/eq/band/4/amp/set 0.5" (there are certainly other ways to approach it too).
If you want to send control signals to multiple parts of the audio player at the same time, you could use a "bundle" composed of a separate message for each parameter you're altering. All the messages will then be handled simultaneously.
Because of this, data does not need to be hierarchical. If you really want to, you can consider the endpoints to be "keys" and the control data to be "values" in a hierachical key -> value map, just to prove that everything you could do with hierarchical data is still possible using OSC. But I would encourage you to not think of it in that abstraction because the fact that OSC is built to pass messages, and not non-contextualized data is a large part of what differentiates it from other protocols like json streaming.
Shameless plug time! I've been using SuperCollider to design a machine-learning solo electroacoustic music performance platform, called Sonic Multiplicities. You can hear it at http://www.sonicmultiplicities.net/.
Also, I find Armstrong's fanaticism grating, and am inherently distrustful of anybody who claims that their community and language has all the good ideas. Because odds are, they're wrong.
You might find this[0] blog post by Armstrong interesting, wherein he comments on some things that were/are done wrong in Erlang. I would also argue that this 'fanaticism' is made to look more severe as you can find at least 4 different talks where Armstrong argues for the same things over and over.
In any case, Armstrong is actually very humble, which I think seems fairly obvious if you actually watch a few talks and read more than a couple of blog posts by him.
Could you point to something specific where you feel Armstrong misrepresents a strong side of Erlang, or is it just a general thing?
(Edit: Just to add, I think it should be said that the Erlang team created their favorite language and that it probably represents the most useful language in use today to them, so I don't think it's surprising that they'd see more good ideas there than in any other language. On top of that Armstrong is a very straight shooter, it seems, so I think this might be something to take into account.)
0 - http://joearms.github.io/2013/05/31/a-week-with-elixir.html
That is if you don't expect to handle more than one connection at a time in a reasonable way. You can certainly do this loop:
server_socket = create_and_listen_on_socket(...)
while True:
csock = server_sock.accept()
csock.send(process(csock.recv()))
csock.close()
TCP setup and teardown is not too cheap (especially compared to sending UDP packets).A receive-only message can be handled in TCP:
server_socket = create_and_listen_on_socket(..., MAXCONN)
while True:
csock = server_sock.accept()
msg = csock.recv()
csock.close()
process(msg)
A UDP loop has similar steps, minus the accept and close. The complexity to use threads or async to run process() is the same for UDP and TCP.For a normal data volume remote command protocol, the simple TCP loop is more than adequate. Just set a reasonable MAXCONN to queue up client connection requests, which can drop connections if there are too many requests, just like UDP dropping packets.
Edit: I don't believe lock can be avoided in Sonic Pi server when handling concurrent incoming commands, whether it's written in Erlang or not. Sonic Pi I believe has a single audio device, which makes it a shared resource. Concurrent access to a shared resource has to be managed with lock somewhere along the call path. In that case a single thread TCP server is perfectly fine, serving as a lock as well.
UDP is unreliable and has small packet size. The TCP code is emulating that simple requirement. Why expand the requirement? Looping to read fully would block on one connection. One rogue client would hold up the whole server.
You are only guaranteed to get the bytes in the correct order but you have to find the boundaries between messages yourself.
... which requires three round trip times for the client and server to agree upon.
UDP gives the flexibility of sending "payload" starting from packet #1. There may be session identifiers in that package, too.
But he was talking about TCP imposes the notion of session on the application while UDP doesn't, which are false. If he meant the TCP connection as a session, why he led the discussion to managing session with locks and threads in application. Then he talked of the ease of Erlang handling session while other languages having a hard time, which was an exaggeration. Session management is a solved problem, in many languages by many people.
Context here:
"I guess I’d underestimated the difficulty of implementing Anything-over-TCP in a sequential language. Just because it’s really really easy in Erlang doesn’t mean to say it’s easy in sequential languages."
Since Erlang isn't a declarative language in the larger sense, that's my guess what he meant.
Also I hope some day one would show how to implement most of functionality of Hadoop orders of magnitude more effeciently in bunch of Plan9 shell scripts. Or in Erlang with 9P protocol driver.
Something like:
$ echo foo | host2:grep bar | host3:grep bazSockets, or ssh? To get a remote shell, you're going to want ssh.
> Joe Armstring: A Badass Way to Connect Programs Together