My student's talk, which is worth watching and has a cool demo of using the same system for bulk facial recognition, is here: https://www.usenix.org/conference/nsdi17/technical-sessions/...
Daniel Horn, Ken Elkabany, Chris Lesniewski, and I used the same idea of finer-granularity parallelism at inconvenient boundaries, and initially some of the same code (a purely functional implementation of the VP8 entropy codec, and a purely functional implementation of the JPEG DC-predicted Huffman coder that can resume encoding from midstream) in our Lepton project at Dropbox to compress 200+ PB of their JPEG files, splitting each transcoded JPEG at an arbitrary byte boundary to match the filesystem blocks: https://www.usenix.org/conference/nsdi17/technical-sessions/...
My students and I also have an upcoming paper at NSDI 2018 that uses the same purely functional codec to do better videoconferencing (compared with Facetime, Hangouts, Skype, and WebRTC with and without VP9-SVC). The idea is that having a purely functional codec lets you explore execution paths without committing to them, so the videoconferencing system can try different quantizers for each frame and look at the resulting compressed sizes (or no quantizer, just skipping the frame even after encoding) until it finds one that matches its estimate of the network's capacity at that moment.
One conclusion is that video codecs (and JPEG codecs...) should support a save/restore interface to give a succinct summary of their internal state that can be thawed out elsewhere to let you resume from that point. We've shown now that if you have that (at least for VP8, but probably also for VP9/H.264/265 as these are very similar) you can do a lot of cool tricks, even trouncing H.264/265-based videoconferencing apps on quality and delay. (Code here: https://github.com/excamera/alfalfa). Mosh (mobile shell) is really about the same thing and is sort of videoconferencing for terminals -- if you have a purely functional ANSI terminal emulator with a save/restore interface that expresses its state succinctly (relative to an arbitrary prior state), you can do the same tricks of syncing the server-to-client state at whatever interval you want and recovering efficiently from dropped updates.
For lambda, it's fun to imagine that every compute-intensive job you might run (video filters or search, machine learning, data visualization, ray tracing) could show the user a button that says, "Do it locally [1 hour]" and a button next to it that says, "Do it in 10,000 cores on lambda, one second each, and you'll pay 25 cents and it will take one second."