History of Apache Storm and Lessons Learned
nathanmarz.com
nathanmarz.com
> That led me to the idea of "spouts" and "bolts" – a spout produces brand new streams, and a bolt takes in streams as input and produces streams as output.
Spout I can grasp; why "bolt"? Why these two, together? I'd have really appreciated an extra sentence or two here as to why he chose these terms -- naming is so crucial, and the fact that both "spouts" and "bolts" sound like sources/producers to me (one of water, another of electricity...?) made it harder to grasp how Storm works.
"Spout" makes sense as something that's a source of a stream. I haven't been able to figure out the metaphor for "bolt" at all, particularly as something that processes water (stream/spout...) in some way.
To be super clear about it, a "topology" in Storm is a directed acyclical graph. The nodes are called "bolts" and the edges are called "spouts".
Not really. Edges are streams and nodes are either spouts or bolts. Spouts have no predecessors and inject external data into the system. Bolts gather one or more input streams and produce zero or more output streams.
These names are a bit confusing, but only during the very first steps of a project. Quickly every developers grasp the concepts.
I have nevertheless a concern regarding spouts and bolts. In practice, when a topology is refactored, bolts have frequently to be transformed into spouts. This occurs when a topology is split in two to ease deployment; and when a persistent queue is added along a stream. This highlights that the distinction between spouts and bolts is bit artificial.
Jeezuz. Call a spade a spade.
An aside: when Nathan started work on Storm, there was a competing streaming platform being worked on inside Yahoo, called "S4". It had many similarities to Storm, and some key differences.
It is still in the incubator: http://incubator.apache.org/s4/ (I have no connection with S4, just know some people who worked on it)
You and Kyle Kingsbury are two "younger" guys that I always enjoy reading whatever you write.
Storm is distributed stream processing.
The two work together very well. Often, a server will publish messages to Kafka (or another queue), which are then read and processed by Storm.
http://www.slideshare.net/gschmutz/kafka-andstromeventproces...
Slides 42 and 43 describe an architecture with all three ... hadoop, storm and kafka. Seems like a beast and I don't have applications but pretty cool combo of technologies.
https://github.com/apache/storm/tree/master/external/storm-k...