I wrote a children's book / illustrated guide to Apache Kafka
gentlydownthe.stream
gentlydownthe.stream
Sorry, it's not really a fair criticism, and I did like the art style - I just don't like Kafka much because I had to deal with it in places where it didn't make much sense.
I've heard of it being used as a sort of message queue for application level events before, but that sounds like a nightmare of trying to reinvent the actor model with 1000x the complexity.
Just like SOA and ESB, the concept isn't the problem, is the technical constraints of the design at the time. Decoupling and messaging isn't bad, but having a legacy message queue on physical hardware doesn't really hold up. Any derived architecture faces the same problem.
Then again, Kafka isn't an actor model implementation, and Akka isn't a partitioned redundant stream processing system, they don't have all that much overlap ;-)
Not saying Akka can replace Kafka but many of the issues around availability, durability and reliability have been attempted to be solved in Akka.
If you set "log.retention.hours = -1" and "log.retention.bytes = -1" in kafka config, kafka stores messages forever.
In a game for example, user-inputs and other events can be produced in Kafka, then reconstruct the entire game-state by reading and processing the kafka-stream from start to finish. It has an advantage over most DBs because it's real-time.
You can use also chronological data streams to represent data structures more complicated than a simple array. For example, a tree can be represented while preserving chronology. This is far from the ideal use case however.
This is called event sourcing.
Kafka itself kind of only solves half the problem, because it doesn't offer indexes as a builtin, so you have to build your indexed state yourself, or maybe even use a "traditional" datastore as a de facto cache. But since you've moved all the hard distributed part of the problem into kafka, that part is not so bad.
It is easy to scale horizontally to massive volumes of data, and any issues in the processing pipeline can be fixed without losing any raw sensor data (restarting the consumers from the last known valid point).
An RSS feed must be powered by something underneath, and any of those tools can do the job. In most situations it would be extremely impractical to use a relational database for this kind of thing. If you are getting thousands of messages in per second, which is not uncommon, no transactional database will give you enough write performance and won't be able to handle too many queries per second for clients polling for updates, like an RSS feed would require. Note that caching queries is almost useless here because the latest content is updated every few milisseconds.
Kafka, as pretty much any other queueing datastore, is optimized for append-only writes with no deletes or updates. Reading from the end of the queue is extremely fast and sequential reads down the stream are quite fast too. Random reads such as the ones commonly handled by SQL databases are either not available or are less efficient than with SQL.
That said, Kafka can be used to pass any kind of message between applications: from simple text messages and small JSON data to vídeo frames, protobuf messages and other types of chunks of serialized data.
It is also a very durable data store for immutable, time-ordered data, and is widely used in the financial world to store transaction logs.
RSS feeds are great as they are easy to implement and debug.
Surprising too. I would have expected a far darker approach to this subject!
Perhaps something like 60's horror comics genre, or better, like those Jack Chick religious tracts where people end up in burning in Hell forever as an immediate consequence of a sin against god.
“Is your orgs politics so complicated that direct team-to-team communication has broken down? Is your business process subject to unannounced violent change? Bogged down by consistent DB schemas and API versioning? Tired of retries on failed messages? Introducing Kafka by Apache: an easy to use, highly configurable, persistently stored ephemeral centralized federated event based messaging data analytics Single Source of Truth as a Service that can scale with any enterprise-sized dysfunction!”
I don't think many children are into Franz Kafka. Kafka is more for cynical grownups. (And teenagers about to become it). It was very Kafkaesque, having to read Kafka in school ..
"Alas", said the mouse, "the whole world is growing smaller every day. At the beginning it was so big that I was afraid, I kept running and running, and I was glad when I saw walls far away to the right and left, but these long walls have narrowed so quickly that I am in the last chamber already, and there in the corner stands the trap that I am running into."
"You only need to change your direction," said the cat, and ate it up.
I have read about 10 of these "children's books about programming" and while I enjoy them myself, I find that they lack most of the things that grab children's attention, such as repetition and visual-only sub-plots.
This is not to criticize the use of an illustrated fable as a storytelling device for adults: They're great! It's just sad that we have to frame an illustrated fable as "for children" in order for it to be accepted. I think it says something about how content dictates and narrows the expected presentation format, sometimes to the detriment of clear communication.
Not my cup of tea (loved the current Kafka one)
I prefer when books for learning use a consistent art style.
The example that I think illustrates my point best is: http://arthur-johnston.com/hacker_writes_a_childrens_book/ (which was posted here in 2017: https://news.ycombinator.com/item?id=15879519) The book works well as entertainment for grown-up programmers: "G is for Garbage Collector/when something's no longer needed/it frees up the memory/so your program is speeded." You can look inside the book on Amazon for more examples. I judge this rhyme as too advanced for any child below 8 that I've read aloud to. Most theories about cognitive development agree with Piaget (1896–1980) in that children have a hard time grasping difficult abstract concepts before the age of 12–13, so there is at least some "scientific backing" to my hunch: https://en.wikipedia.org/wiki/Cognitive_development#Concrete... (This is only a hunch though; the children I've read aloud to are all picked from the same group – children of family and friends.)
My observation is that children seem to need _lots_ of concrete verbal and visual imagery in order to stay interested. You can also get away with more abstract themes if the actual text is lyrically well-crafted, like Dr. Seuss' books (or André Bjerke's children's rhymes in Norwegian, my mother tongue). The most successful attempts at teaching programming to children that I've seen, seem to give the children a lot of practical tasks they can work on (e.g. the Hello Ruby book series). You also have some programming elements in Minecraft that children could pick up, because the concepts are implemented as concrete objects.
All of this makes me suspect that the main audience for these "children's books" teaching programming is adults that already know a fair bit about the subject. And that's completely all right by me, because it gives the books traction and brings the book's readers entertainment.
PS: You also have the Javascript/HTML/CSS for babies series, which are only jocularly presented as children's books: https://imgur.com/eOYc8fC
I think he likes it mostly because I tell him that’s what I do at work.
The guy above you mentioned manga-guides to stuff, which utterly fail at their job, which again, is to distill key-information in an entertaining, easy-to-read, general (but fully accurate) manner.
Note: I'm not against creative attempts to explain technical concepts. But the form to me seems odd, and that it feels like we're producing very short tutorials in a childrens format. That's even weirder.
These kinds of explanations tend to focus on the most critical/important concepts, and help validate (or dispel) assumptions I've made about the tech.
This focus on analogy also lets the author tell the story faster, because I already understand:
- What an otter is
- That rivers flow
- The water flowing down a river that forks will be spread across those forks
- etc...
Depending on the strength of the analogy, it's possible to get the reader on the same page much faster than an intro/tutorial that must first explain foundational concepts just to get to the basics of the technology itself.
The art and animation in this is great, but I feel like the author's talents are wasted on a document with no audience. Make a kids' book instead!
But I can't deny they are very popular!
I'm going to try and fit this into as many sentences as I can get away with!
There's no shortage of dry technical documentation, so seeing something akin to outsider art in that space was really refreshing.
Personally I would love to see more technical books come with a soundtrack!
Different people, different preferences I guess.
This prompted me to make an awesome repo to collect children's books on technical topics (which I have seen a few on HN).
https://github.com/searchableguy/awesome-illustrated-guides
I couldn't find a similar list. If there's one, let me know.
Google Machine Learning - https://cloud.google.com/products/ai/ml-comic-1
Google Federated Learning - https://federated.withgoogle.com/
>> "When a mommy and a daddy love each other very much, the daddy wants to give the mommy a special gift.
>> So he buys a "stay-at-home" server."
P.S. This sentence is hard to grasp for me "This Unawareness helps Decouple systems that produce events from otters that read events." https://www.gentlydownthe.stream/#/20
https://medium.com/hackernoon/intuitive-rl-intro-to-advantag...
It simply means that the producer doesn't need to maintain a list of listeners. It just throws the event into the stream, and assumes that anyone who wants to read it will be able to do so.
RSS vs an email list, I guess.
More generally there's this list of similar fiction https://fiftysevendegreesofrad.github.io/hard-comp-fi-fictio...
Is there a typo at https://www.gentlydownthe.stream/#/22 ?
Current version: "First, they dropped large stones into the river, splitting each topic into a smaller number of streams, or Partitions."
I know nothing about Kafka, but I think maybe this should be: "First, they dropped large stones into the river, splitting each topic into a number of smaller streams, or Partitions."
(my emphasis for both versions) "
Not necessarily reminiscent of a children's book, but still a more entertaining way to learn a programming language than most guides.
Even if you do need to keep your events long term, why not use something like Eventstore?
It feels like whenever Kafka developers had to make any tradeoffs they chose performance over everything else. I would only ever use it again if I the sheer data volume makes it impossible to use an alternative.
As far as I know, Redis Streams does not offer all of that on its own. You could certainly build a lot of that for Redis yourself (and someone probably has already), but then you start getting back that complexity.
So if you have high requirements for performance, durability and availability, are don't mind the added complexity then Kafka is worth a look.
edit: the differences in consumer groups confuse me as well. With Redis streams, if you have a single stream with a single consumer group with multiple consumers, each consumer will get a new message at different times so processing of each message may happen out of order. That seems... fine to me. Apparently that is not the case with Kafka which makes me ask, why have multiple consumers in a single group if one will be blocked on the other?
Redis doesn't have partitions, and instead has a single host managing a stream, while taking on all the consumer tracking functionality. It's fast, but not as scalable, and eventually if you keep growing then scale is how you get speed.
You can compromise by using hashing or other logic to group related messages into the same partition. For example, hashing by user-id lets you process events for a particular user in order, while processing all events by users in parallel. If you can divide into logically consistent but globally isolated boundaries then this setup works very well.
And this is life after Kafka: https://www.gentlydownthe.stream/#/24
https://www.gentlydownthe.stream/#/11: “In the rivers gleam” doesn’t seem quite right; I can’t decide whether that should be “In the river’s gleam” (the river doing the gleaming) or whether it was intended to be “In the rivers gleam” (the events doing the gleaming, in multiple rivers) in which case I suspect it should be a singular river.
Now I want a Docker children's book please.
The container ecosystem is pretty complicated :-).
[1] https://kubernetes.io/docs/setup/production-environment/cont...
Kafka has its flaws, but it really served us well. We have Python Data Engineers who focus on distributed system design[1], and Kafka is one of the team's least finicky open source components, but it is used everywhere, and it basically enables the entire rest of the real-time data processing stack.
If you want a deep dive into event streaming systems (logs) that also gives some examples of big data systems they interact with, I highly recommend "The Log: What every software engineer should know about real-time data's unifying abstraction," from LinkedIn, the creators of Kafka.
Their engineering blog also has many other interesting articles on their data systems.
[0] https://engineering.linkedin.com/distributed-systems/log-wha...
on page 12, i think you mean "streams" rather than "seams". no? also, the page title is "gently down the steam", is an R missing?
That said, I feel like the otters could have just made a bulletin board to solve their problem
[1]: https://www.amazon.com/SCADA-Me-Book-Children-Management/dp/...
also mathematic and science in ancient india were taught using easy to remember poems, for example Pascal's triangle:
https://lvnaga.wordpress.com/2014/10/21/meru-prastarah-or-pa...
also different numeric series represented as poems https://en.m.wikipedia.org/wiki/Katapayadi_system
we definitely need more of these that teach complex mathematics and topics like machine learning.
And it can't be aimed at children, given that amount of technical jargon (we suddenly go from otters and rivers and bees to headers, keys, values and timestamps). And, well, children don't need a book about Kafka.
But then I realized that that's what makes it a realistic story ;)
Jokes aside, the best explanation I've ever seen. It should be a must-do for people who want a quick and understandable introduction to the concepts.
Otherwise quite well done!
I have to say, i was looking for a slide on how the otters cleverly know what messages they've already read, and that was unfortunate, but kudos!
Quick feedback: My nephew prefers page turns (curling) on their iPad. I assume that he prefers it over a powerpoint style slides that we currently have on your book. Well, at least he spent some time flipping through the animated pages!
Perhaps they usually work with streams and couldn’t resist applying a familiar solution to the problem?...
It's often hard to find a legitimate use case for Kafka/RabbitMq/etc. where a much simpler solution won't do.
Likely the `expires` header is the problem (it's a date in the past)
The survey asked "would you be interested in hardcover", my answer is no. Wouldn't know where to put it. Also less paper the better.
The fact that since January 2000 nearly 16k people "found this helpful" opened my eyes to how Amazon book reviews really do have the potential to be "classic".
Thank you for sharing this and really brightening my day.
Why not just make a recommendation.
My homestay family sister Anna, 9 years old, told me that I spelled "sed" wrong :D
ed likes to write books. His home is near the C.
ed has short black hair on his head, with a curl.
ed has a cat with a long tail. His cat likes to read books. He likes it when you tcl his tummy.
Last time, ed looked into the clear sky. It's a nice view. On top of the arch, he saw a bird.
"Whatis your name?" asked ed.
"Awk" said the bird.
"Can you talk to a man?" asked ed.
"Yes" said awk. "I read books. But only sum words."
"Can you find all the words that start with a B?" asked ed.
"books" "black" "bird" "but" "bin" "big" "bits" "bash" sed awk.
"That's right! You're very good at reading. Which place do you live in?" asked ed.
"Now I live in a bin." said awk.
Awk is hungry.
To goto the field, he crawls through a pipe.
He gets some food from the field, and puts it on the table to eat it with a fork.
Yummy! He takes a byte.
Oh no! Now he's too fat to go back home.
"Don't worry", says ed. "I can sort that out."
"My cat will cut food from the field and keep on sending it through the pipe for you."
"Now I need to go to the toilet." says awk.
They go to the toilet. Wait while IP, says ed.
Awk sits down and unzips his trousers.
awk does one poo, two, three, four, five, six, seven, eight!
The wc has 8 bits of poo. There's a big login the heap. Stinky!
Flush! The poo goes away through another pipe.
The awk must wash his hands. But the sink is grotty and it leaks.
ed has strong arms. He can flex his muscles. "I'll make it better with tar!" he says.
With a bash and a clang, the sink is flowing again.
"Eat less next time." says ed. "The only animal who eats more than you is the GNU."
"I was hungry!" says awk. "What did you expect?"
It's time to leave. There's no reason to stay here idle. "Welcome back to our group any time. cu later!"
Questions:
1. Who does ed meet in the story? (a) Gnu (b) Awk (c) Python
2. Can you remember a word starting with B? (any of "books" "black" "bird" "but" "bin" "big" "bit" "bash")
3. What takes food to the awk, and flushes the poo away? (a) Pipe (b) Bin (c) Field
4. What do you do after going to the toilet? (a) Watch YouTube (b) Brush your teeth (c) Wash your hands
"Then, the otters would decide which part of the river to put the message in"
I'd imagine a child looking at this going like "why are they doing this to the river". Why are they are throwing things into the river. :D
I think Kafka needs the equivalent of OpenAPI and Redoc, a simple spec and document generator, but for groups of consumers and producers rather than single applications. This would increase the tractability of complex systems, but it would also let you see the system getting more complex over time, even when you haven't reached the pain point yet.
(I view OpenAPI as a near-failure, but, good luck to everyone trying.)
https://www.asyncapi.com/docs/specifications/v2.0.0#referenc...
id Identifier Identifier of the application the AsyncAPI document is defining.
This format is fine, but the tool I'm looking for is something that will read the definitions of multiple AsyncAPI documents (multiple applications) and show how their inputs and outputs connect, so I can answer a question like, "When application X publishes message Y on channel Z, which applications consume that message?"AsyncAPI gets me 90% there by defining the service spec and providing code to parse it. It's possible somebody has already written the rest; I'll have to see.
[0] https://www.asyncapi.com/docs/specifications/v2.0.0#fixed-fi...
Your question makes as much sense as: why is <new database X> any different from SQL?
OK. So I didn't get what Kafka is.
With REST APIs (first few pages of the book), services talk directly with each other, with event sourcing (the rest of the book) services talk with an event store (Kafka in the book) as the intermediary.