MillWheel: Fault-Tolerant Stream Processing At Internet Scale [pdf]
db.disi.unitn.eu
db.disi.unitn.eu
- https://blog.twitter.com/2013/streaming-mapreduce-with-summi...
- http://samza.incubator.apache.org/
Of the three I'm most interested in Summingbird. Unlike Google's system I can actually download it, and it seems to provide a better query abstraction than Samza. I haven't spent much time investigating any of these systems, so I might be incorrect in this assessment.
Edit: Guess I was thinking about spark[1] but it doesn't really fit.. others?